Skip to content
unmatrixai

On-prem deployment

On-prem open-source LLMs for regulated industries

On-prem LLM deployment means running open-source language models such as Llama, Mistral, Qwen or gpt-oss on hardware inside your data centre or private cloud, so no prompt or document leaves your network. Unmatrix AI deploys and operates these models for banks, insurers, healthcare providers, utilities and public bodies whose agentic workflows cannot use hosted APIs.

Last updated

When is on-prem the right call?

  • Data residency or sovereignty rules that forbid third-party processing
  • PHI, cardholder data or classified material that cannot leave your boundary
  • Air-gapped networks with no outbound internet
  • Predictable cost at high volume, where per-token pricing becomes the largest line item
  • Contractual control over model versions so behaviour does not change under you

Will an open model be good enough?

Open models have closed most of the gap with hosted frontier models on enterprise tasks, and the gap that remains is task-specific. We do not guess. During scoping we run your evaluation set against candidate open models and a frontier model, and show you the difference before you buy hardware.

  • The same eval suite on every candidate model
  • Task-level results, not benchmark scores
  • Routing so the hardest cases can use a larger model if your policy allows

The stack

What do we deploy?

Models
Llama, Mistral, Qwen, gpt-oss and DeepSeek-class open-weight models, chosen per task and licence. Smaller models for extraction and classification, the largest you can serve for reasoning.
Serving
vLLM or Hugging Face Text Generation Inference behind an OpenAI-compatible API, with batching, quantisation and tensor parallelism tuned to your GPUs.
Gateway
A model gateway that routes requests, enforces budgets and logs usage, so agents are written once and can move between models.
Embeddings and rerankers
Open-source embedding and reranking models served the same way, so retrieval-augmented generation runs entirely inside.
Platform
Kubernetes or your existing orchestration, with GPU scheduling, autoscaling and blue-green model upgrades.
Guardrails
The same PII and PHI detection, prompt injection screening, evals and audit logging as our cloud deployments, running on-prem.

Hardware

What hardware does it need?

GPU classes
NVIDIA H100 or H200 for the largest models, A100 or L40S for mid-sized, and a single L4 is enough for small extraction models.
Quantisation
8-bit and 4-bit quantisation cut memory by two to four times with a measurable, usually small, quality cost that we verify on your evals.
Sizing
We size from your measured workload: tokens per request, requests per minute, latency target. A pilot on rented GPUs de-risks the purchase.
Air-gapped delivery
Model weights, containers and dependencies are delivered as a signed bundle for networks with no internet access.

Operations

Who runs it?

  • Monitoring of latency, throughput, GPU utilisation and quality drift
  • Model upgrades tested against your evals before rollout, with rollback
  • Security patching of the serving stack
  • Capacity reviews as agent volume grows
  • Handover to your platform team with runbooks, or a support retainer

Questions

What platform teams ask

Which open-source LLMs do you recommend?

It depends on the task, the licence and your GPUs. The Llama, Qwen, Mistral and gpt-oss families cover most enterprise needs. We test candidates on your own evals rather than recommend from a leaderboard.

Is on-prem more expensive than API pricing?

At low volume, yes. At sustained high volume the hardware pays back and the cost is predictable. We model both against your projected usage before you decide.

Can we run this in a private cloud instead of our own data centre?

Yes. A private VPC with dedicated GPU instances gives the same data boundary. Many regulated clients start there.

Do agents built for hosted models work on open models?

Yes, through the gateway. Prompts may need tuning and the evals tell us where. Tool calling is now standard in the major open model families.

Can it be fully air-gapped?

Yes. We deliver signed bundles and operate without outbound access. Updates arrive the same way.

Who operates it after handover?

Your platform team, with our runbooks, or us under a retainer. Either way it lives in your environment.

Next step

Want to know if an open model is good enough for your workflow?

Bring a sample of real cases. We will run them against two open models and a frontier model and show you the numbers before you buy a GPU.