Skip to job description
Home Security Fabric AI Servers FAQ Resources Careers Contact Sales

Platform Engineering

Principal Platform & Release Engineer

Build the installable, signed, observable, upgradeable, and recoverable enterprise release path.

Remote/Hybrid, United States; periodic production-site and customer travel Full-time Reports to Chief Technology Officer

Job Description

We’re looking for a Principal Platform & Release Engineer - someone who understands that enterprise software delivery is part of the product, not an administrative step after development.

Think of this as the person responsible for turning the Woven Security Fabric and Agentic Orchestration Governance platform into an installable, versioned, signed, observable, upgradeable, recoverable, and supportable enterprise system.

You don’t stop when the code works on a developer’s machine. You build the path by which a clean Summit Series system can receive an approved release, verify it, install it, operate it, update it, diagnose it, roll it back, and recover from failure.

You will work across application engineering, infrastructure, containers, networking, identity, security, model serving, observability, testing, packaging, and documentation. Your work will remove dependency on tribal knowledge and create the production discipline required for Fortune 50 deployments.

This is a senior individual-contributor role. You will design architecture, write code, build automation, diagnose failures, document systems, and directly own production outcomes.

In Your Day-to-Day, You Will

  • Own the software release lifecycle for the Woven Security Fabric and Agentic Orchestration Governance platform.
  • Build reproducible release processes so the same source, configuration, dependencies, and build instructions produce verifiable software artifacts.
  • Design and maintain installation automation for Summit Series hardware and supported customer-controlled environments.
  • Create versioned, signed, and verifiable release artifacts, including integrity checks, dependency records, configuration manifests, and software bills of materials.
  • Build infrastructure-as-code and configuration-management systems that establish consistent production environments without undocumented manual steps.
  • Develop installation paths for connected, restricted-network, and air-gapped deployments, including dependency mirroring, offline artifact distribution, update media, and verification procedures.
  • Design and maintain continuous integration and release pipelines covering compilation, packaging, unit testing, integration testing, security checks, deployment validation, and release approval.
  • Build automated test environments capable of validating the platform against clean hardware, supported configurations, and expected customer deployment patterns.
  • Own upgrade, migration, rollback, backup, restoration, and disaster-recovery procedures, and test those procedures under realistic failure conditions.
  • Create observable systems, including appropriate logs, metrics, traces, health checks, alerts, performance indicators, and operator dashboards.
  • Develop secure support bundles and diagnostic tools that help Island Mountain troubleshoot systems without unnecessarily exposing customer information.
  • Establish configuration standards for identity, networking, certificates, secrets, model services, storage, logging, agent services, and customer-specific integrations.
  • Prevent customer-specific forks by turning repeated implementation needs into controlled, documented configuration options and supported extension points.
  • Improve software supply-chain security, including dependency governance, artifact signing, secret management, vulnerability remediation, provenance, and controlled release authorization.
  • Design load, performance, endurance, failover, and recovery tests that show how the platform behaves under operational stress.
  • Create release notes, operator guides, deployment runbooks, troubleshooting procedures, compatibility records, and known-issue documentation for every production release.
  • Partner with the Hardware Production Engineer to validate drivers, firmware, accelerator support, operating-system images, performance assumptions, and recovery procedures on production Summit Series systems.
  • Partner with the Agentic Security & Governance Engineer to make security-policy tests, authorization tests, audit tests, and control validation part of the release process.
  • Partner with the Deputy Forward-Deployed Enterprise Engineer to turn field lessons into supported product improvements rather than one-time customer workarounds.
  • Participate directly in early customer deployments, incident resolution, release decisions, and technical demonstrations.

Requirements

  • Significant experience in platform engineering, release engineering, site reliability engineering, infrastructure engineering, backend systems, or a closely related field. This is commonly gained through eight or more years of work, but equivalent evidence of production ownership is welcome.
  • Strong Linux systems experience, including services, processes, permissions, networking, storage, certificates, logging, packaging, and system diagnosis.
  • Experience building or operating production deployment systems using containers, orchestration platforms, virtual machines, system services, or comparable technologies.
  • Strong experience with continuous integration, automated testing, artifact management, release pipelines, and infrastructure as code.
  • Ability to write production-quality code or automation in at least one language such as Go, Python, Rust, TypeScript, Java, or a comparable language.
  • Experience designing installation, upgrade, rollback, backup, restoration, and migration procedures.
  • Strong understanding of networking fundamentals, TLS, certificate management, DNS, proxies, service discovery, firewalls, and restricted egress.
  • Experience with secure software delivery, dependency management, secrets, artifact integrity, access control, and release authorization.
  • Ability to diagnose failures across application code, operating systems, containers, storage, networking, identity systems, model infrastructure, and hardware.
  • Experience building observability and operational-diagnostics systems using logs, metrics, traces, health checks, alerts, and support tooling.
  • Ability to create precise technical documentation that another engineer can execute without relying on verbal explanation.
  • Ability to make sound production decisions during incidents, deadlines, pilot deployments, and incomplete-information conditions.
  • High agency, strong technical judgment, and willingness to work directly in the codebase and deployment environment.

A specific college degree is not required. Evidence of systems you have personally deployed, operated, upgraded, recovered, or repaired carries more weight than pedigree.

Preferred Experience

  • AI or machine-learning infrastructure, model serving, GPU workloads, inference systems, vector databases, or distributed agent platforms.
  • On-premise enterprise software, private-cloud appliances, edge systems, or customer-controlled infrastructure.
  • Air-gapped or restricted-network deployment and software-update processes.
  • Enterprise identity integration, including OIDC, OAuth, SAML, LDAP, Active Directory, certificate authorities, or workload identity.
  • Kubernetes, Nomad, Docker, Podman, systemd, Terraform, Ansible, Nix, Bazel, Helm, or comparable technologies.
  • Artifact signing, software provenance, dependency attestations, or software bill-of-materials systems.
  • Self-hosted registries, package mirrors, offline dependency management, and disconnected release distribution.
  • High-availability architecture, stateful distributed systems, backup systems, and disaster-recovery exercises.
  • Enterprise security reviews, technical due diligence, security questionnaires, or regulated customer deployments.
  • Early-stage product environments where the platform engineer helped create the operating discipline rather than inheriting it.

What Success Looks Like

During your first 30 days, you will map the current architecture, deployment path, dependencies, release process, configuration model, testing gaps, and recovery risks.

Within 60 days, you will establish a versioned release process, clean-environment installation workflow, automated integration test path, release documentation standard, and initial rollback procedure.

Within 90 days, another qualified engineer should be able to take a clean Summit Series system, install an approved release from documented artifacts, validate the installation, upgrade it, roll it back, collect diagnostics, and restore it without relying on undocumented founder knowledge.

Within six months, Island Mountain should have a predictable release cadence, controlled compatibility matrix, automated security and regression testing, supported air-gapped update process, tested recovery procedures, and operational telemetry suitable for enterprise deployment.

About the Team

We’re Island Mountain, building secure, customer-controlled infrastructure for enterprise AI and autonomous systems.

Our product suite includes Summit Series private compute, the Woven Security Fabric, and Agentic Orchestration Governance - technology designed to give organizations greater control over where AI workloads run, which tools and information agents can access, how actions are authorized, how resources are consumed, and how system activity is recorded.

We are moving from advanced product development into repeatable production and enterprise deployment.

We work quickly, but we do not confuse speed with disorder. We document decisions, automate repeated work, test what we claim, and build systems that can survive outside the development environment.

The platform you build will become the operating foundation for every customer deployment that follows.

Ready to Apply?

Send the following to [email protected]:

  • The title of the position
  • Your résumé, LinkedIn profile, or GitHub profile
  • Examples of platforms, deployment systems, or release processes you personally built
  • A brief account of a difficult production failure or recovery you owned
  • Three to five points describing what you would aim to accomplish during your first 90 days

The door is open.

---

Apply directly

Show us the work you have owned.

Applications are delivered to [email protected]. Include concrete evidence: systems built, failures diagnosed, controls tested, or deployments carried through.

PDF, DOC, or DOCX.

Or apply by email

How we hire

Demonstrated ability over pedigree.

Island Mountain values demonstrated ability, sound judgment, direct ownership, and clear communication.

We encourage applications from people whose experience was gained through traditional employment, independent work, military or public service, entrepreneurship, open-source contribution, skilled technical practice, or nontraditional education.

You do not need to meet every preferred qualification to apply. Requirements describe the work that must be performed; preferred experience identifies backgrounds that may help someone become effective more quickly.

Our hiring process may include:

  1. An introductory conversation focused on motivation, experience, and role alignment
  2. A technical conversation with the responsible executive or engineering lead
  3. A practical work sample based on a realistic Island Mountain problem
  4. A final discussion about ownership, operating style, compensation, and first-90-day expectations

We do not use irrelevant puzzles or performative interview exercises. Our practical assessments are designed to show how you think, communicate, prioritize, document, and execute.

Island Mountain is hiring. The door is open.