Skip to content
unmatrixai

Agent architecture

RAG agents with hybrid search

A RAG (retrieval-augmented generation) agent answers questions by first retrieving the relevant passages from your documents and then generating a response grounded in those passages, with citations. Unmatrix AI builds RAG agents with hybrid search, combining keyword and vector retrieval with reranking, so answers hold up on exact terms such as policy numbers as well as on meaning.

Last updated

RAG with hybrid searchDocuments are ingested with their permissions, parsed and chunked, then indexed twice: a keyword index and a vector index. A question retrieves from both, results are merged, filtered by permission and reranked, and the model answers only from those passages with citations.INGESTIONDOCUMENTSwith access groupsPARSE & CHUNKby document structureKEYWORD INDEXBM25VECTOR INDEXembeddingsANSWERINGQUESTIONin the user's wordsHYBRID RETRIEVEboth indexesRERANKpermission filter firstANSWERcitations or not found
RAG with hybrid search

When is this the right architecture?

  • Staff spend time searching policies, contracts, manuals or tickets for answers that already exist
  • Answers must cite the source and respect who is allowed to see it
  • The corpus changes and the answers must change with it, without retraining anything
  • Exact identifiers matter as much as concepts: clause numbers, part codes, drug names

When is it the wrong one?

  • Tasks that require acting in systems rather than answering. Pair it with a workflow or autonomous agent
  • Questions whose answer is not written down anywhere
  • Corpora with no access-control story

Mechanics

How does it work?

  1. 01

    Ingest

    Connectors pull documents from SharePoint, Confluence, Google Drive, Box or a file share, keep their permissions, and watch for changes.

  2. 02

    Parse and chunk

    PDFs, decks, spreadsheets and scans are converted to text with layout preserved. Chunks follow the document's structure (sections, clauses, tables), not a fixed character count.

  3. 03

    Index twice

    Each chunk gets a keyword (BM25) index entry and a vector embedding. Metadata such as document type, owner, date and access group is stored alongside.

  4. 04

    Retrieve with hybrid search

    A query runs against both indexes. Results are merged and reranked by a cross-encoder. Permission filters apply before ranking, so a user never sees a passage they could not open.

  5. 05

    Generate with citations

    The model answers only from the retrieved passages and cites them. If the passages do not support an answer, it says so instead of guessing.

  6. 06

    Learn from corrections

    Operators can flag wrong or missing answers. Those become eval cases and, where needed, ingestion fixes.

Controls

What guardrails ship with it?

  • Permission-aware retrieval inherited from the source system
  • Citations on every answer with a link to the passage
  • An explicit not-found path instead of a fabricated answer
  • Prompt-injection screening of ingested documents, because documents can carry instructions
  • PII and PHI detection at ingestion with redaction where policy requires
  • Retrieval and answer evals against a held-out question set

Worked example

A typical build: policy assistant for frontline staff

A contact centre works from hundreds of policy documents that change monthly. Agents put customers on hold to find the right paragraph.

  1. 01

    Ingest

    Policy documents from SharePoint are ingested nightly with their access groups. Superseded versions are retired automatically.

  2. 02

    Ask

    An agent types the customer's question in their own words: can a tenant claim for water damage from the flat above?

  3. 03

    Retrieve

    Keyword search finds the clause containing the phrase escape of water. Vector search finds the related exclusions. The reranker orders them.

  4. 04

    Answer

    The assistant gives a three-line answer with two citations and a link to the clause. If the policy version matters, it says which one it used.

Handling time drops and answers become consistent across the team. Wrong answers are traceable to a passage, which is how the corpus gets fixed.

Stack

What do we typically build it with?

Retrieval
Elasticsearch or OpenSearch for BM25, pgvector, Qdrant or Weaviate for vectors, a cross-encoder reranker
Parsing
Layout-aware PDF and Office parsing, OCR for scans, table extraction
Embeddings and models
Hosted or open-source embedding models, Claude or GPT-class models for generation, open-source models on-prem where required
Sources
SharePoint, Confluence, Google Drive, Box, ServiceNow, file shares, data warehouses
Evals
Retrieval recall and answer faithfulness measured on a held-out set, in CI

Questions

What people ask about rag agents

What is hybrid search and why does it matter?

Hybrid search runs a keyword search and a vector search on the same query and merges the results. Keyword search is exact on identifiers and rare terms. Vector search matches meaning when the wording differs. Enterprise questions need both.

How do you stop the agent from making things up?

It can only answer from retrieved passages, it must cite them, and it is evaluated on faithfulness. When the passages do not contain the answer it says so and offers the closest documents.

Will people see documents they are not allowed to see?

No. Access groups are captured at ingestion and applied as a filter before ranking. If permissions change in the source, the next sync applies them.

Is RAG the same as fine-tuning?

No. RAG keeps knowledge in documents that can change daily and be cited. Fine-tuning changes the model's behaviour and is rarely the right tool for facts.

How large a corpus can it handle?

Millions of chunks is routine. The constraints are parsing quality and permission modelling, not index size.

Can a RAG agent run entirely on-prem?

Yes. Parsing, indexes, embeddings and the generation model can all run inside your network on open-source components.

Next step

Pick one workflow. We will be at your office in two weeks.

A thirty-minute call to find the right first workflow, followed by a written scoping note. No deck, no pilot.