Job Description
We’re looking for a Lead Hardware Production & Systems Integration Engineer - someone who can take a secure AI system from an advanced prototype to repeatable production.
Think of this as the person responsible for turning the Summit Series into a controlled, buildable, testable, traceable, and supportable hardware product.
You don’t just assemble servers. You qualify components, resolve supply-chain constraints, harden firmware, document system configurations, automate production testing, and make certain that every unit leaving Island Mountain performs to the same approved standard.
This role sits at the intersection of hardware engineering, secure systems integration, production operations, Linux infrastructure, supply-chain management, and field deployment. You should be equally comfortable reviewing a bill of materials, diagnosing a GPU or network failure, configuring firmware, writing a validation script, working with a supplier, and explaining production risk to company leadership.
The physical platform is the foundation beneath the Woven Security Fabric and Agentic Orchestration Governance platform. You will be accountable for making that foundation repeatable and trustworthy.
In Your Day-to-Day, You Will
- Own the Summit Series reference architecture from prototype through pilot production, including compute, accelerators, memory, storage, networking, chassis, power, cooling, firmware, and peripheral requirements.
- Create and maintain the production bill of materials, approved vendor list, qualified component alternates, lifecycle status, lead-time assumptions, and per-unit cost models.
- Qualify hardware components and configurations through performance, compatibility, thermal, power, endurance, and failure testing.
- Build and refine production systems directly, working hands-on with rack-mounted servers, GPU platforms, storage arrays, networking equipment, power systems, management controllers, and secure hardware components.
- Establish secure firmware and hardware configuration standards, including BIOS or UEFI configuration, secure boot, TPM use, management-controller hardening, firmware version control, device identity, and restricted administrative access.
- Design the secure provisioning process that transforms approved hardware into a customer-ready Island Mountain system.
- Develop automated production acceptance tests covering component health, memory, storage, network throughput, accelerator performance, thermal stability, firmware state, system identity, and software readiness.
- Create burn-in and stress-testing procedures capable of identifying unstable components before systems reach a customer environment.
- Establish component and configuration traceability, including serial-number records, firmware versions, production dates, test results, customer assignment, configuration history, and chain-of-custody documentation.
- Build clear assembly, provisioning, inspection, packaging, shipping, installation, and return procedures that another qualified technician can follow without relying on undocumented founder knowledge.
- Manage technical relationships with suppliers, distributors, original equipment manufacturers, and contract production partners, including component availability, quality concerns, warranty issues, substitutions, and production schedules.
- Develop production capacity plans for small pilot batches and future volume, including labor requirements, equipment needs, long-lead purchases, spare inventory, and supplier risk.
- Create field-service and repair procedures, including diagnostics, replacement standards, spare-parts planning, warranty handling, and return-material authorization workflows.
- Support customer deployments remotely and on site, particularly during early pilots where hardware, networking, power, cooling, and software installation must work as one system.
- Partner closely with the Platform & Release Engineer to make sure software images, installers, drivers, firmware, system dependencies, and recovery procedures are validated against actual production hardware.
- Partner with the Agentic Security & Governance Engineer to ensure hardware configuration supports device identity, trusted execution requirements, auditability, and the larger Island Mountain control model.
- Measure and improve the environmental performance of the system, including energy use, cooling requirements, component longevity, repairability, utilization, and responsible end-of-life planning.
- Own the production record, documenting defects, corrective actions, configuration changes, supplier issues, and lessons learned from every system built.
Requirements
- Significant hands-on experience in hardware systems engineering, production integration, server engineering, data-center infrastructure, or a related field. This experience is commonly gained through seven or more years of work, but demonstrated capability matters more than a specific timeline.
- Direct experience building, configuring, testing, or supporting high-performance compute systems, GPU servers, private-cloud appliances, edge-computing platforms, or comparable infrastructure.
- Strong working knowledge of server components, accelerators, memory, storage, networking, power, cooling, firmware, and system-management interfaces.
- Strong Linux systems knowledge, including installation, drivers, hardware diagnostics, logging, networking, storage, permissions, and automation.
- Ability to write practical automation and diagnostic scripts using tools such as Python, Bash, PowerShell, or comparable languages.
- Experience working with BIOS, UEFI, secure boot, TPMs, baseboard management controllers, firmware updates, and remote-management interfaces.
- Experience creating or maintaining bills of materials, component qualification records, build instructions, test procedures, or other controlled production documentation.
- Ability to diagnose complex failures across hardware, firmware, drivers, operating systems, networking, and application layers.
- Understanding of supply-chain risk, component substitution, lifecycle management, vendor qualification, and failure analysis.
- Ability to make sound engineering decisions under time, availability, and cost constraints without lowering the product’s security or reliability standard.
- Clear written and verbal communication, including the ability to document technical work and explain production risks to both technical and nontechnical leaders.
- High agency and independence, with the discipline to establish repeatable processes rather than repeatedly solving the same problem by hand.
- Ability to travel periodically to production locations, suppliers, demonstrations, and customer deployment sites.
A specific college degree is not required. We care more about the systems you have personally built, tested, repaired, documented, and shipped.
Preferred Experience
- GPU infrastructure, AI inference systems, high-performance computing, or model-serving appliances.
- Private-cloud, on-premise, restricted-network, or air-gapped systems.
- Original equipment manufacturer, original design manufacturer, contract manufacturing, or systems-integrator environments.
- PXE provisioning, Redfish, IPMI, infrastructure automation, image management, or fleet-management systems.
- Hardware-rooted identity, cryptographic attestation, secure boot chains, hardware security modules, or trusted platform architecture.
- Production test engineering, failure-mode analysis, quality-control systems, or corrective-action processes.
- Data-center power planning, thermal analysis, rack integration, network architecture, and deployment-site readiness.
- Enterprise hardware warranty, field support, repair, and return programs.
- Security-sensitive, regulated, critical-infrastructure, or government customer environments.
What Success Looks Like
During your first 30 days, you will assess the existing Summit Series architecture, component dependencies, documentation, supply-chain exposure, production process, and testing gaps.
Within 60 days, you will establish a controlled reference bill of materials, secure configuration baseline, assembly process, provisioning workflow, and initial production acceptance test.
Within 90 days, Island Mountain should be able to build at least two equivalent systems from the same documentation, provision them through the same process, test them against the same acceptance standard, and produce a complete configuration record for each unit.
Within six months, Island Mountain should have a working pilot-production process, qualified component alternatives, measured defect data, documented field-service procedures, and a costed production plan for increasing volume.
About the Team
We’re Island Mountain, building secure, customer-controlled infrastructure for enterprise AI and autonomous systems.
Our product suite includes Summit Series private compute, the Woven Security Fabric, and Agentic Orchestration Governance - technology designed to give organizations greater control over where AI workloads run, which tools and information agents can access, how actions are authorized, how resources are consumed, and how system activity is recorded.
We are moving from advanced product development into repeatable production and enterprise deployment.
We operate as a small, senior team. Titles matter less than ownership, evidence, and shipped work. The people joining at this stage will directly influence the product, the production standard, and the operating culture of the company.
Ready to Apply?
Send the following to [email protected]:
- The title of the position
- Your résumé or LinkedIn profile
- A portfolio, project history, GitHub profile, or examples of systems you personally built
- A brief explanation of the most difficult hardware or infrastructure failure you have diagnosed
- Three to five points describing what you would aim to accomplish during your first 90 days
The door is open.
---