Home Forward Deployed Security Fabric FAQ Resources Blog About Start a Scoping Call
Answers · Human to Human

Straight answers about sovereign AI.

What a deployment looks like, what it protects, what it takes to run. Swappable hardware. Swappable open-weight models. No cloud. Every spec and every quote, human to human.
Air-Gappable Quote-Based
Air-Gapped AIHIPAAAttorney-Client PrivilegeITAR / CMMCTribal SovereigntyWoven Security FabricZero Standing PrivilegeAny Open-Weight ModelDeepSeek V4-FlashQuote-Based
Start Here

Two halves of one sovereign stack.

We install the configuration and the governance together: on-premise hardware running open-weight models entirely on your own floor, and the Woven Security & Governance Fabric sitting inline on every action the AI takes. The questions below cover both, plus compliance and how quotes work. There is no public price list, because sizing this is a human conversation and always has been.

The Configuration

Whatever the Work Needs

Any hardware, any open-weight model, any data server rack, air-gappable throughout. A couple of desktop-class machines and a few laptops, or a large estate on rack-mounted accelerators. Discovery decides which.

See how a configuration gets decided →
The Trust Plane

Woven Security & Governance

Identity, guardrails, approval gates, and an audit ledger inline on every tool call, with zero standing privilege. Woven into every unit, and deployable cloud-side or fully air-gapped.

Explore the Security Fabric →
Sovereign AI, Explained

The basics, in plain terms.

What is local AI inference hardware, and how is it different from cloud AI?

Local AI inference hardware is a physical server with GPUs that runs AI models on your premises instead of sending your data to someone else's data center. "Inference" means the system processes your prompts and generates responses from pre-trained models. It does not train new models; it runs them locally.

With a cloud AI service (ChatGPT, Claude, Gemini, Azure OpenAI), your prompts leave your network, get processed on the provider's GPUs, and the response comes back over the internet. The provider sees your data and processes it on shared infrastructure. With local hardware the entire cycle happens inside your building: prompt goes to your GPU, response comes back from your GPU, nothing leaves your network.

You own the hardware, you own the open-weight models, and you control every variable. And because every Island Mountain unit carries the Woven Security & Governance Fabric, governance travels with the work instead of living in a vendor's cloud. The tradeoff is honest: cloud gives you the latest proprietary models without buying hardware; local gives you data sovereignty, no recurring token fees, and models you control, while open-source releases trail proprietary ones by weeks to months.

Does sovereign AI hardware work fully air-gapped, with no internet?

Yes. Every Island Mountain system is designed to operate completely air-gapped. The models, the inference engine (vLLM and Ollama), and the interface (OpenWebUI) all run locally. Once the system is set up and models are loaded, you can disconnect the ethernet cable entirely and everything still works.

Users reach the system through a web browser on your local network. As long as the device and the server share a network, even an isolated one with no internet, inference runs normally. No phone-home. No license checks. No cloud dependencies. The only things that ever need internet are downloading new models and applying OS security updates, both optional and on your schedule.

Island Mountain applies a documented air-gap configuration before delivery, disabling every outbound feature at the environment-variable level, and confirms it before shipping. Your IT team can independently verify every setting on the delivered system.

What does a forward deployment look like?

It starts with us on-site, watching your people work, before anything gets specced. We sit with the staff who run the workflow: the senior paralegal who knows which intake questions matter, the records clerk who can tell you which form gets filled out wrong every time, the compliance officer who knows where the real risk lives. We learn the work the way they run it, not the way the SOP describes it, and fill in the gaps in conversation.

That study becomes a system design. Only then does hardware enter the picture, sized to what we found. After the build we hand it off and train your team, because a deployment that leaves behind people who understand the thing is the one still running a year later. The full method is on the Forward Deployed page.

How long are you on-site, and who do you need access to?

Long enough to understand the work, which depends entirely on how tangled it is. A single-department workflow is a short engagement. An organization with a dozen handoffs between departments takes longer, and we would rather tell you that than quote you a tidy number that turns out to be fiction.

Who we need is more important than how long. Not a steering committee: the people who do the job, for enough hours that we see the exceptions and not just the happy path. Those hours pay for themselves several times over, and they are the first thing organizations try to cut.

Who should NOT buy sovereign AI hardware?

We would rather tell you now than after a serious capital purchase. If your organization needs the absolute latest proprietary models (GPT-4o, Claude Opus, Gemini Ultra) on release day, local hardware will not deliver that; open-source closes the gap over time, but there is always a gap.

If your data is not sensitive and cloud AI is working fine for a handful of casual users, the sovereignty story may not justify owning infrastructure. And if your organization has no IT capability at all and cannot rack or maintain a server, know that after the included 30-day setup support your team owns the day-to-day, though ongoing support retainers are available.

Where local hardware earns its place is exactly where cloud creates exposure: regulated data, privileged data, sovereign data, and agentic workflows that must be governed and audited. If that is you, the rest of this page is written for you.

The Security Fabric

Woven security & governance.

The model can be tricked. The fabric is where the tricked action is stopped and logged. Here is how the governance layer works, whether it runs cloud-side or fully air-gapped.

What is the Woven Security & Governance Fabric?

The Woven Security & Governance Fabric (WSF), built as Lamprey, is the trust and control plane that sits inline on every action an AI agent takes. Identity, guardrails, approval gates, and an audit ledger on every tool call. It gives operators cloud visibility, data classification, compliance evidence, and approval-gated remediation without handing a third party standing access or your operating context.

On Island Mountain hardware it is woven into every unit, so findings, samples, credentials, and audit trails stay inside your boundary. It can also run on its own, cloud-side or air-gapped, for teams that need governance across an estate they already operate. For the full platform, see the Security Fabric page.

Can I run the fabric cloud-side, or does it have to be air-gapped?

Both. It runs two ways: outside your organization through the cloud, or inside as an air-gapped deployment. Either way, findings, samples, credentials, and audit trails stay in your custody, aligned to your existing security and governance model.

The architecture is hub-and-gateway: lightweight gateways run near the resources they inspect while evidence, status, and workflow are centralized where operators can act. It covers AWS, Azure, GCP, Kubernetes, and databases, runs inside your environment with no scan-data SaaS, and keeps access scoped through ephemeral credentials and audited approvals.

How does the fabric stop agentic attacks like prompt injection and exfiltration?

Agentic attacks ride the same rails as legitimate work, so the controls live outside the model, inline on every call. Prompt injection, agent-jacking, tool poisoning, data exfiltration, memory poisoning, and rogue agency all try to turn ordinary tool use against you.

Orchestration flows down through runtime and tools into the fabric, where identity, guardrails, approval gates, and the audit ledger evaluate each action before it executes. Sensitive actions are approval-gated, access is scoped and short-lived, and every request and approval is recorded. The model can be fooled; the fabric is where the fooled action is caught, stopped, and logged. The homepage security map shows the flow end to end.

What is "zero standing privilege"?

Zero standing privilege means nothing holds broad, permanent access. Scan access is minted just in time, scoped to the operation, and revoked automatically. Write access and remediation are granted only after approval and stay fully audited.

We treat least privilege as a product constraint, not a dashboard label: a scanner or an agent must earn each action instead of accumulating trust it can later be tricked into misusing. That reduces the blast radius of both honest mistakes and active compromise.

Does the fabric produce compliance evidence and audit trails?

Yes. It preserves who requested, approved, and executed each sensitive action, connects technical findings to control families, and keeps audit-ready context so remediation reads as an operational control rather than a black box.

Evidence collection aligns to control mapping and repeatable review workflows that fit inside an existing compliance program, and it stays in your custody instead of becoming another external data exhaust. That evidence is what turns "we ran AI locally" into something you can show an auditor.

Compliance & Sovereignty

The regulated-data questions.

We are not your compliance attorney, and nothing here is legal advice. What local hardware does is remove the third-party processing vector; what the fabric adds is the evidence. Your organization still owns the operational program.

Does running AI locally comply with HIPAA?

Running AI on local hardware eliminates the data-transmission vector that creates HIPAA exposure in the first place. With a cloud service, Protected Health Information (PHI) leaves your network and is processed on shared infrastructure controlled by a third party. Under 45 CFR §164.402, that transmission can constitute a reportable disclosure event. With local inference, PHI never leaves your premises, and there is no business associate to sign a BAA with because none is involved.

Local hardware does not by itself make your entire workflow HIPAA-compliant. You still need access controls, audit logging, encryption at rest, and staff training. The hardware removes the cloud transmission risk, the fabric supplies the audit trail, and your organization owns the rest. Island Mountain is not a compliance attorney and does not certify HIPAA compliance. See our Medical Practices page for how this fits a clinical program.

Can a law firm use local AI without waiving attorney-client privilege?

The core issue with cloud AI for law firms is this: client data sent to a third-party API is data disclosed to a third party. Under ABA Model Rule 1.6, attorneys have a duty to make reasonable efforts to prevent unauthorized disclosure of client information, and courts have found privilege waived when confidential information is processed through systems outside the firm's control.

Local AI keeps every prompt, document, and response on a server inside your office, behind your firewall, on hardware you own. Your firm still needs internal policies on who accesses the system, what data goes in, and how outputs are treated in work product, but the fundamental architectural risk of third-party disclosure is removed. See our Law Firms page for more.

Can a tribal government run AI without cloud providers?

Yes. It is one of the primary use cases Island Mountain hardware is built for. Tribal nations exercise inherent sovereignty over constituent data, enrollment records, health information, and governance documents. Routing that data through cloud AI means processing sovereign data through jurisdictions and infrastructure outside tribal authority.

Local AI keeps everything on tribal premises, under tribal law, processed by hardware the nation owns outright. No data leaves the reservation, no cloud terms of service apply, and no federal or state jurisdiction touches the processing. The pre-installed models carry permissive open-source licenses, so the nation owns them with no usage restrictions, and the system runs fully air-gapped if your posture requires it. See our Tribal Nations page for OCAP-aligned workflows.

Does local AI hardware meet ITAR and CMMC requirements for defense contractors?

ITAR restricts processing and storage of controlled technical data to environments that prevent foreign access. Cloud infrastructure with multinational operations, overseas data centers, or foreign-national employees creates structural compliance risk. Hardware you own, operate on US soil, and control with US-person-only access removes the cloud processing vector entirely.

The open-weight models are not themselves ITAR-controlled; your data is the controlled element, and local hardware keeps it off infrastructure you do not control. Island Mountain does not provide ITAR or CMMC certification. That compliance covers physical security, personnel screening, access controls, and documentation well beyond the hardware, and the Woven Security & Governance Fabric supplies the audit evidence that supports it. See our Defense Contractors page for CUI, CMMC, and DFARS workflows.

Is DeepSeek V4-Flash safe to use with sensitive organizational data?

The safety question usually comes from two concerns: the model's origin and data handling. Here are the facts.

DeepSeek V4-Flash is an open-weights model released under the MIT license. When you run it on Island Mountain hardware, it executes entirely on your local GPUs. No data is sent to DeepSeek, to any Chinese server, or to any external endpoint. The weights are publicly available and independently audited by the open-source community. The model does not phone home and cannot: it is a file on your drive being processed by your GPUs.

The risk with DeepSeek exists only when you use its cloud API, which routes your data through their servers. That is not how Island Mountain systems work; ours run the model locally, air-gapped if you want, with zero external communication. Treat it like any other software asset: verify the download hash, apply your acceptable-use policy, and run it on infrastructure you control.

The On-Premise Hardware

Hardware, models, and how it scales.

What hardware will we end up with?

Whatever Discovery says the work needs, and there's no catalog to pick it from. No set specs, no reference build, no standing recommendation, because the recommendation doesn't exist until we've seen what you do. Then we recommend the stack purchase and we deploy and configure it together, in your building.

Any hardware, any open-weight model, any data server rack. A small office doing real work on a couple of desktop-class machines and a few laptops is one honest answer. A large estate on rack-mounted accelerators with adjacent data server racks is another. Neither one is the upsell, and there's no ladder between them that anybody gets walked up.

It's your building and your budget, so if you want something other than what we'd put in, we'll install that instead. How a configuration gets decided walks through what has to be on the table before anyone names hardware at all.

Can you size it before you've seen how we work?

No, and that's deliberate. Which models you'd run, how deep the context has to go, and how many people are hitting it at once all move the answer, and none of those are knowable from a phone call. Anyone who sizes your system on the first call is sizing their inventory.

Guessing it up front is the most common way on-premises AI deployments fail. The box arrives, it's wrong for the work, and nobody involved wants to be the one to say so. The sizing falls out of Discovery, after we've watched the work happen. That's slower, and it's the reason these land.

Which models do we get, and are we stuck with them?

Open-source, open-weight models chosen to fit your work, and no, you are never stuck with them. The model layer is as swappable as the hardware. A deployment typically lands with a reasoning and general-purpose model like DeepSeek V4-Flash at FP8, a strong generalist like Llama 4 Scout, and a third picked for your domain: DeepSeek R1 70B Distill (MIT, explicit chain-of-thought, useful for legal and compliance work) or Qwen 3 72B (Apache 2.0, enterprise general performance).

All carry permissive open-source licenses with no usage restrictions. You own them outright and switch between them in the OpenWebUI dropdown. Any compatible open-weight model that fits your VRAM can be pulled through Ollama or OpenWebUI, so as the frontier moves you swap rather than renegotiate. That is the point of open weights: no lab holds a meter on you, and nobody can deprecate the model your workflow depends on. For the first 30 days we walk your team through model management directly.

What is OpenWebUI, and how does it work?

OpenWebUI is a free, open-source, browser-based interface for local AI models. Think of it as a private version of ChatGPT that runs entirely on your network. No cloud account, no external data transmission.

You open a browser on any device on your network, navigate to the server's local address, and start prompting. It includes a dropdown to switch between installed models, full conversation history stored locally on the server, and an admin panel for user management and per-user model access. Conversations live on the server's drive, not in any cloud.

OpenWebUI needs no command-line knowledge. If someone can use ChatGPT, they can use OpenWebUI, and it is pre-configured on every system before it ships.

How small can this be, and how big does it get?

Small enough for an office with no server room, and big enough to feed several departments at once. Both ends are real deployments, neither one is the upsell, and where you land on that spread comes out of Discovery rather than a configurator.

What doesn't change anywhere on it: it's a one-time hardware purchase with no per-seat fees and no token metering, it stays on your premises, and it gets built with room to grow so you add capacity as the work grows instead of replacing what you already own. How a configuration gets decided covers what that turns on.

What does it need from our facility?

Power, cooling, network, and somewhere it can physically live, and all four follow from the configuration. A couple of desktop-class machines want a normal wall outlet and the air conditioning you already have. Rack-mounted accelerators want a dedicated high-amperage circuit, real server-room airflow, and a rack to sit in.

Which is why we scope it with your facilities people before anything is ordered, rather than after it turns up on a loading dock. Facilities gets a veto and it is a great deal cheaper to hear it early. If a circuit needs installing, a licensed electrician can do it, and we'll tell you exactly what to ask for.

How can OpenWebUI be air-gapped if it is a web application?

The "web" in OpenWebUI refers to the browser-based interface, not internet dependency. It is a self-hosted application that runs entirely on the server. Users reach it through a browser pointed at the server's local address, the way you reach a router's admin panel. No cloud account, no external API calls for inference.

Out of the box it does include outbound features: model downloads, update checks, HuggingFace Hub access, community sharing, web search for RAG, and telemetry. None are required for inference, and Island Mountain disables every one before shipping. The key environment variables (OFFLINE_MODE, HF_HUB_OFFLINE, ENABLE_COMMUNITY_SHARING, ANONYMIZED_TELEMETRY, ENABLE_RAG_WEB_SEARCH, SAFE_MODE) are set to their air-gapped states during the build and confirmed before shipping.

Your IT team can verify it independently by inspecting the container environment, or by watching a monitored network segment for zero outbound connections during operation.

I already own DGX Sparks or a Strix Halo box. Will Island Mountain work with hardware you didn't sell?

Yes. We wire the serving stack, stand up the software, and train your team on hardware from any vendor, including NVIDIA DGX Spark units, AMD Ryzen AI Max (Strix Halo) machines, and servers another integrator shipped and never configured. Since we don't have a catalog to protect, gear we didn't sell isn't a problem for us the way it is for a shop attached to a product.

The gap that kills most on-premises AI projects is systems design, not hardware: the serving framework, the quantization choices, the multi-user layer, the air-gap configuration, which is the same work we put into every deployment.

Small unified-memory boxes are real local AI hardware with honest limits: strong for one to three users on mixture-of-experts models, weak on concurrency and long-document prompt processing. Our DGX Spark vs RTX PRO 6000 comparison runs the numbers. If the box you own fits your workload, we will get it serving; if it does not, we will say so before you spend more. Talk to the builder.

Quotes, Delivery & Trust

Owning it, human to human.

There is no public price list and no self-serve checkout. Every spec, every number, and every commitment comes from a real conversation with the person who builds the hardware.

How does pricing work?

Pricing is scoped to your workload and headcount, and every number comes from a conversation, not a configurator. You tell us what you are running, how many people need access, and which models matter; we spec the right system and put a written quote in front of you.

Each build is a one-time hardware purchase with no subscription fees, no per-token charges, and no recurring software licensing. The models are open-source, OpenWebUI is open-source, and the operating system is Ubuntu Server LTS. There is no public price list because sizing sovereign infrastructure is a human decision, not a shopping-cart line item. Talk to the builder to get a number, or call 1-341-441-8740.

How does an engagement start, and how long until we are running?

It starts with a scoping call, then time on-site with your people, then a written quote scoped to what we found. We do not spec speculatively. Once the design is agreed, that triggers sourcing of your specific GPUs and components through authorized NVIDIA channels with full procurement documentation. Terms are handled directly, human to human, no automated checkout.

Hardware lead time runs roughly 3 to 5 weeks, longer when the build calls for parts sourced per order. That clock overlaps the workflow study rather than following it, so the schedule is usually set by how quickly we can get the right hours with the right people, not by shipping. The system is configured and air-gapped before it lands, so once it is racked and powered (208V/30A) most organizations run their first real prompts the same day. Every deployment includes 30 days of hands-on support with direct access to the engineer who built it.

What warranty comes with it, and what if a component fails?

Every system ships with a 1-year hardware warranty covering all components, including GPUs, CPU, RAM, storage drives, and power supplies. If a component fails within the first year, we handle the replacement.

GPU failures are managed through supplier RMA agreements with documented replacement timelines, and we hold a 20% warranty reserve per unit so replacements are not waiting on budget. You contact the builder directly (phone 1-341-441-8740 or email), we diagnose remotely, and we ship the part or arrange a swap. Extended warranty options are available at the time of purchase, and after warranty, support is available per incident or on an annual retainer.

You are a new company. Why trust a first purchase, and what if you go out of business?

This is a fair question, and you should ask it. Island Mountain is founded by Basho Parks, who hand-builds each rack. The company is new; the team is not.

Whatever a deployment lands on, it's standard enterprise hardware bought through authorized channels with established supplier RMA agreements, backed by a 1-year warranty and a 20% warranty reserve per unit. Nothing exotic, nothing single-sourced, nothing you'd have trouble replacing. The software is entirely open-source, so nothing is proprietary and nothing is locked.

If Island Mountain closed, your system would keep running. Ubuntu Server LTS, vLLM, Ollama, OpenWebUI, and open-weight models depend on none of our servers. You would lose our support and warranty, but the system needs no license, no activation, and no connection to us. OS updates come from Ubuntu, drivers from NVIDIA, models from public libraries, and any competent Linux administrator can maintain it. Vendor independence is exactly the problem we are solving.

Can local AI keep up with cloud AI as models improve?

Not on day one, but within weeks to months, yes. Cloud providers ship proprietary models the moment they control the infrastructure. Open-source models lag that cycle, so when a new flagship launches you will not have an equivalent on your local hardware that afternoon.

What has happened consistently is that open-weight models close the gap fast. DeepSeek V4-Flash, Llama 4 Scout, and DeepSeek R1 all reached or exceeded the proprietary models they followed within weeks to months. Your system can run any new open-source model that fits its VRAM: you download it through Ollama or OpenWebUI, with no hardware change. And because every deployment is built with room to grow, you add capacity as your needs grow instead of starting over.

Why pay for a build when I could build a similar server myself for less?

You can, and if you are a developer running experiments, you probably should. The guides exist and the software stack is the same: Ollama, vLLM, OpenWebUI, and open-weight models like DeepSeek V4-Flash and Llama 4 Scout. With the staff to build, configure, test, and maintain it, a DIY box will run inference.

Here is what a DIY build does not include. Professional procurement documentation: every accelerator in a deployment has a documented chain, purchase receipts from authorized channels, a serial-number registry, and RMA history, which consumer cards from Amazon or Newegg do not carry. Professional RMA chains that keep replacement parts inside the same documentation trail, which matters for HIPAA, ITAR/DFARS, and CMMC audits. Direct builder support: you talk to the person who assembled your system, not a ticket queue.

The honest answer: if you are experimenting, build your own. If you are a regulated organization putting AI into production where provenance and evidence matter, that is what the difference covers.

Does Island Mountain have a written doctrine?

Yes, and it's public. The Island Mountain Doctrine is the company's point of view written down before we knocked on anyone's door: the chain of command (technology serves people, people serve missions, missions serve communities, communities preserve civilization), the seven-phase method behind every engagement, the security posture, and the 28 principles a deployment answers to. It's versioned like an engineering document, revised in the open, and silent edits never happen.

Read the full text: The Island Mountain Doctrine.

Summary: Island Mountain configures and deploys air-gappable, on-premises AI inference on any hardware and software combination the work calls for, each deployment carrying the Woven Security & Governance Fabric that governs and audits every action an AI agent takes. There is no catalog, no set hardware specification, and no standing recommendation: the configuration is chosen per customer after an on-site immersion and Discovery, then deployed and configured together. Systems run frontier open-weight models fully offline, suited to organizations bound by HIPAA, ITAR/CMMC, attorney-client privilege, tribal data sovereignty (OCAP), and FERPA. There is no public price list: every quote is scoped human to human.
Still Have Questions?

Ask a human.

One conversation, no sales pitch. Tell us about your organization and what you are trying to accomplish with sovereign AI, and we will give you a straight answer, and a real number.

Or call directly: 1-341-441-8740

Prefer to read first? Browse the resource library or start from the platform overview.

Explore the Platform

Keep exploring

Answers are a starting point. Here is the deeper reading, and the fastest way to reach a human.

FAQ

Straight answers on deployment, model selection, air-gapping, governance, and exactly what runs on a on-premise rack.

Read the FAQ →

Resources

Direct-answer briefs on air-gapped inference, HIPAA, ITAR/CMMC, the CLOUD Act, and on-premises AI cost.

Browse resources →

Contact

Tell us about your organization and what you're trying to accomplish. One builder, one phone call, real answers.

Contact sales →