Skip to content
unmatrixai

Security and guardrails

Guardrails, PII and PHI protection, and prompt injection defence

Unmatrix AI treats personal data (PII) and health data (PHI) as a design constraint rather than a compliance checkbox, and builds guardrails and prompt injection defence into every agent we ship. This page describes the controls, where they sit in the system, and how we test them.

Last updated

Personal and health data

How do we handle PII and PHI?

Personal and health data is detected before a model sees it, minimised to what a step needs, kept inside your boundary, and retained on your schedule.

Detect before the model sees it
Inputs and documents pass through PII and PHI detection (names, identifiers, contact details, health information) before any model call. Where policy requires, values are redacted or replaced with tokens the agent can carry but not read.
Minimise
Agents receive the fields a step needs, not the whole record. Tool responses are trimmed by the same rule.
Keep it in your boundary
Models run in your cloud account or on-prem. Prompts, documents and traces are stored in your systems with your encryption keys. Nothing is used to train models.
Retention you set
Traces and logs follow your retention schedule. Redacted traces can be kept for audit while raw inputs are dropped.
Health data
For PHI we work inside your HIPAA programme: business associate agreements with model providers where hosted models are used, or open-source models on-prem where they are not.
Access
Agents inherit least-privilege roles from your identity provider (Okta, Microsoft Entra). Every read and write is logged with who, what and why.

Prompt injection

How do we defend against prompt injection?

Prompt injection is text an agent reads, from a user, a document or a web page, that tries to change what the agent does. The defence is layered: keep instructions and data apart, screen what comes in, limit what any single run can do, and check what goes out.

Separate instructions from data
System instructions, user input and retrieved content are kept in distinct, labelled parts of the prompt. Content from documents, emails and web pages is treated as data that may contain instructions, never as instructions.
Screen inputs and retrieved content
A classifier flags instruction-like text in inputs and ingested documents. Flagged content is quarantined or shown to a person before the agent acts on it.
Constrain what an attack could do
Tool allowlists, least-privilege credentials, step limits and spend limits mean a successful injection can only do what the agent could already do, and never write to a system of record without a person.
Validate outputs
Outputs are checked against schemas and policies before they reach a system or a customer. Free text does not flow into actions.
Sandbox execution
Any code or tool execution runs in an isolated environment with no network access beyond the allowlist.
Human approval for side effects
Sending, paying, changing or deleting waits for a person when the action is above an agreed threshold.
Red-team before go-live
We test each system against the OWASP Top 10 for LLM Applications, including direct and indirect prompt injection, data exfiltration and excessive agency, and keep the cases in the eval suite.

Guardrails

What guardrails ship with every system?

  • Evals built from your real cases, run on every prompt, model or code change
  • Human review queues for actions above an agreed risk threshold
  • An audit log of every read, decision and write, with model version, exportable to your SIEM
  • Permissions per tool inherited from your identity provider
  • Private deployment in your VPC or on-prem, with your keys
  • Model routing with budgets and fallbacks per workflow

Testing

How do we test it?

Evals from real cases
A held-out set of real inputs, including the awkward ones, scored on every change.
Adversarial set
Injection attempts, exfiltration prompts and out-of-scope requests, replayed in CI. A regression blocks the deploy.
Permission tests
We verify that a user cannot retrieve a passage or trigger an action they could not perform in the source system.
Change control
Every prompt, model or tool change runs the full suite before it reaches production, with a rollback path.

Questions

What security teams ask

Do you send our data to OpenAI or Anthropic?

Only if you choose hosted models, and then through your own enterprise agreement with zero data retention and, for PHI, a business associate agreement. Otherwise we run open-source models inside your network.

What is prompt injection?

Prompt injection is when text an agent reads, from a user, a document or a web page, contains instructions that try to change the agent's behaviour, such as leaking data or taking an unauthorised action. Indirect injection through documents is the version that matters most for enterprise agents.

Can you guarantee an agent will never be manipulated?

No one can. What we can guarantee is that a manipulated agent cannot do more than its tool allowlist and permissions allow, cannot take a risky action without a person, and leaves a trace that shows what happened.

Are you HIPAA compliant?

We build inside your HIPAA programme rather than claiming a certification ourselves: PHI detection and redaction, business associate agreements with any hosted model provider, encryption and audit logs, or a fully on-prem deployment. Your compliance team signs off the design.

How do you handle GDPR?

Data minimisation, retention schedules, EU data residency through your cloud region or on-prem, and the ability to find and delete a data subject's records in traces and indexes.

Do you keep our data to train models?

No. Nothing is used to train any model, ours or a provider's.

Next step

Bring your security team to the first call.

We would rather answer the hard questions in week zero than in week ten. Thirty minutes, and a written design your CISO can mark up.