On-prem deployment
On-prem open-source LLMs for regulated industries
On-prem LLM deployment means running open-source language models such as Llama, Mistral, Qwen or gpt-oss on hardware inside your data centre or private cloud, so no prompt or document leaves your network. Unmatrix AI deploys and operates these models for banks, insurers, healthcare providers, utilities and public bodies whose agentic workflows cannot use hosted APIs.
Last updated
When is on-prem the right call?
- Data residency or sovereignty rules that forbid third-party processing
- PHI, cardholder data or classified material that cannot leave your boundary
- Air-gapped networks with no outbound internet
- Predictable cost at high volume, where per-token pricing becomes the largest line item
- Contractual control over model versions so behaviour does not change under you
Will an open model be good enough?
Open models have closed most of the gap with hosted frontier models on enterprise tasks, and the gap that remains is task-specific. We do not guess. During scoping we run your evaluation set against candidate open models and a frontier model, and show you the difference before you buy hardware.
- The same eval suite on every candidate model
- Task-level results, not benchmark scores
- Routing so the hardest cases can use a larger model if your policy allows
The stack
What do we deploy?
- Models
- Llama, Mistral, Qwen, gpt-oss and DeepSeek-class open-weight models, chosen per task and licence. Smaller models for extraction and classification, the largest you can serve for reasoning.
- Serving
- vLLM or Hugging Face Text Generation Inference behind an OpenAI-compatible API, with batching, quantisation and tensor parallelism tuned to your GPUs.
- Gateway
- A model gateway that routes requests, enforces budgets and logs usage, so agents are written once and can move between models.
- Embeddings and rerankers
- Open-source embedding and reranking models served the same way, so retrieval-augmented generation runs entirely inside.
- Platform
- Kubernetes or your existing orchestration, with GPU scheduling, autoscaling and blue-green model upgrades.
- Guardrails
- The same PII and PHI detection, prompt injection screening, evals and audit logging as our cloud deployments, running on-prem.
Hardware
What hardware does it need?
- GPU classes
- NVIDIA H100 or H200 for the largest models, A100 or L40S for mid-sized, and a single L4 is enough for small extraction models.
- Quantisation
- 8-bit and 4-bit quantisation cut memory by two to four times with a measurable, usually small, quality cost that we verify on your evals.
- Sizing
- We size from your measured workload: tokens per request, requests per minute, latency target. A pilot on rented GPUs de-risks the purchase.
- Air-gapped delivery
- Model weights, containers and dependencies are delivered as a signed bundle for networks with no internet access.
Operations
Who runs it?
- Monitoring of latency, throughput, GPU utilisation and quality drift
- Model upgrades tested against your evals before rollout, with rollback
- Security patching of the serving stack
- Capacity reviews as agent volume grows
- Handover to your platform team with runbooks, or a support retainer
Questions
What platform teams ask
Which open-source LLMs do you recommend?
It depends on the task, the licence and your GPUs. The Llama, Qwen, Mistral and gpt-oss families cover most enterprise needs. We test candidates on your own evals rather than recommend from a leaderboard.
Is on-prem more expensive than API pricing?
At low volume, yes. At sustained high volume the hardware pays back and the cost is predictable. We model both against your projected usage before you decide.
Can we run this in a private cloud instead of our own data centre?
Yes. A private VPC with dedicated GPU instances gives the same data boundary. Many regulated clients start there.
Do agents built for hosted models work on open models?
Yes, through the gateway. Prompts may need tuning and the evals tell us where. Tool calling is now standard in the major open model families.
Can it be fully air-gapped?
Yes. We deliver signed bundles and operate without outbound access. Updates arrive the same way.
Who operates it after handover?
Your platform team, with our runbooks, or us under a retainer. Either way it lives in your environment.
Related
- Read
RAG agents
Document ingestion with permissions, hybrid keyword and vector retrieval with reranking, and grounded answers with citations or an honest not-found.
- Read
Security and guardrails
How every Unmatrix AI system handles PII and PHI, defends against prompt injection, limits what an agent can do, and is tested before go-live.
- Read
Forward deployed
Two to four engineers embed at your office for twelve weeks to take one agentic workflow from scoping to production, then hand it to your team.
Next step
Want to know if an open model is good enough for your workflow?
Bring a sample of real cases. We will run them against two open models and a frontier model and show you the numbers before you buy a GPU.