Agent architecture
RAG agents with hybrid search
A RAG (retrieval-augmented generation) agent answers questions by first retrieving the relevant passages from your documents and then generating a response grounded in those passages, with citations. Unmatrix AI builds RAG agents with hybrid search, combining keyword and vector retrieval with reranking, so answers hold up on exact terms such as policy numbers as well as on meaning.
Last updated
When is this the right architecture?
- Staff spend time searching policies, contracts, manuals or tickets for answers that already exist
- Answers must cite the source and respect who is allowed to see it
- The corpus changes and the answers must change with it, without retraining anything
- Exact identifiers matter as much as concepts: clause numbers, part codes, drug names
When is it the wrong one?
- Tasks that require acting in systems rather than answering. Pair it with a workflow or autonomous agent
- Questions whose answer is not written down anywhere
- Corpora with no access-control story
Mechanics
How does it work?
- 01
Ingest
Connectors pull documents from SharePoint, Confluence, Google Drive, Box or a file share, keep their permissions, and watch for changes.
- 02
Parse and chunk
PDFs, decks, spreadsheets and scans are converted to text with layout preserved. Chunks follow the document's structure (sections, clauses, tables), not a fixed character count.
- 03
Index twice
Each chunk gets a keyword (BM25) index entry and a vector embedding. Metadata such as document type, owner, date and access group is stored alongside.
- 04
Retrieve with hybrid search
A query runs against both indexes. Results are merged and reranked by a cross-encoder. Permission filters apply before ranking, so a user never sees a passage they could not open.
- 05
Generate with citations
The model answers only from the retrieved passages and cites them. If the passages do not support an answer, it says so instead of guessing.
- 06
Learn from corrections
Operators can flag wrong or missing answers. Those become eval cases and, where needed, ingestion fixes.
Controls
What guardrails ship with it?
- Permission-aware retrieval inherited from the source system
- Citations on every answer with a link to the passage
- An explicit not-found path instead of a fabricated answer
- Prompt-injection screening of ingested documents, because documents can carry instructions
- PII and PHI detection at ingestion with redaction where policy requires
- Retrieval and answer evals against a held-out question set
Worked example
A typical build: policy assistant for frontline staff
A contact centre works from hundreds of policy documents that change monthly. Agents put customers on hold to find the right paragraph.
- 01
Ingest
Policy documents from SharePoint are ingested nightly with their access groups. Superseded versions are retired automatically.
- 02
Ask
An agent types the customer's question in their own words: can a tenant claim for water damage from the flat above?
- 03
Retrieve
Keyword search finds the clause containing the phrase escape of water. Vector search finds the related exclusions. The reranker orders them.
- 04
Answer
The assistant gives a three-line answer with two citations and a link to the clause. If the policy version matters, it says which one it used.
Handling time drops and answers become consistent across the team. Wrong answers are traceable to a passage, which is how the corpus gets fixed.
Stack
What do we typically build it with?
- Retrieval
- Elasticsearch or OpenSearch for BM25, pgvector, Qdrant or Weaviate for vectors, a cross-encoder reranker
- Parsing
- Layout-aware PDF and Office parsing, OCR for scans, table extraction
- Embeddings and models
- Hosted or open-source embedding models, Claude or GPT-class models for generation, open-source models on-prem where required
- Sources
- SharePoint, Confluence, Google Drive, Box, ServiceNow, file shares, data warehouses
- Evals
- Retrieval recall and answer faithfulness measured on a held-out set, in CI
Questions
What people ask about rag agents
What is hybrid search and why does it matter?
Hybrid search runs a keyword search and a vector search on the same query and merges the results. Keyword search is exact on identifiers and rare terms. Vector search matches meaning when the wording differs. Enterprise questions need both.
How do you stop the agent from making things up?
It can only answer from retrieved passages, it must cite them, and it is evaluated on faithfulness. When the passages do not contain the answer it says so and offers the closest documents.
Will people see documents they are not allowed to see?
No. Access groups are captured at ingestion and applied as a filter before ranking. If permissions change in the source, the next sync applies them.
Is RAG the same as fine-tuning?
No. RAG keeps knowledge in documents that can change daily and be cited. Fine-tuning changes the model's behaviour and is rarely the right tool for facts.
How large a corpus can it handle?
Millions of chunks is routine. The constraints are parsing quality and permission modelling, not index size.
Can a RAG agent run entirely on-prem?
Yes. Parsing, indexes, embeddings and the generation model can all run inside your network on open-source components.
Related
- Read
Autonomous agents (ReAct)
Agents that reason, call a tool, read the result and decide the next step in a loop, inside a harness that enforces tool allowlists, step limits and human review.
- Read
LangGraph workflow agents
Graph-based workflows where the model decides only at chosen nodes, with checkpointed state and human-in-the-loop interrupts. Example: a resume screening agent.
- Read
On-prem LLMs
Deploying and operating open-source models such as Llama, Mistral and Qwen inside your data centre or private cloud for agentic workflows that cannot use hosted APIs.
Next step
Pick one workflow. We will be at your office in two weeks.
A thirty-minute call to find the right first workflow, followed by a written scoping note. No deck, no pilot.