Agent architecture
Autonomous ReAct agents for enterprise workflows
An autonomous agent on a ReAct architecture is a system in which a language model reasons about a task, calls a tool, reads the result and decides what to do next, repeating that loop until it can produce an output. Unmatrix AI builds these agents for enterprise work where the path to the answer is not known in advance, such as investigating an exception across several systems.
Last updated
When is this the right architecture?
- The route to the answer depends on what each lookup returns
- The work spans several systems and the order of calls varies case by case
- A person doing the job today opens five tabs and follows the evidence
- Each case is valuable enough to justify a few seconds and a few model calls
When is it the wrong one?
- Fixed processes where the steps never change. A LangGraph workflow is cheaper and more predictable
- Answering questions from documents. A RAG agent is the right tool
- Actions with irreversible side effects and no human checkpoint
- Latency budgets under a second
Mechanics
How does it work?
- 01
Task and tools are defined
The agent receives a goal, a set of typed tools (search a system, read a record, draft a message) and the rules it must follow. Tools are the only way it can touch your systems.
- 02
Reason
The model writes a short plan for the next step given everything it has seen so far.
- 03
Act
The harness executes the chosen tool call with least-privilege credentials and a timeout. The model never calls systems directly.
- 04
Observe
The tool result is appended to the agent's working context, including errors and empty results.
- 05
Decide
The model chooses whether to call another tool, ask a person, or finish. Step limits, budgets and the stop condition are enforced by the harness, not the model.
- 06
Output and audit
The final output is validated against a schema, risky actions are routed to human review, and every step is written to the audit log.
Controls
What guardrails ship with it?
- Tool allowlists per agent and per role, inherited from your identity provider
- Hard step and cost limits enforced by the harness
- Human approval for any action that writes to a system of record or contacts a customer
- Schema-validated outputs so downstream systems never receive free text
- Evals replayed from real cases on every prompt, model or tool change
- A full trace of reasoning, tool calls and results for every run
Worked example
A typical build: reconciliation-break investigation
A finance team receives hundreds of unmatched transactions a day. Each one needs an analyst to look in three systems and decide what happened.
- 01
Trigger
A break lands in the queue with an amount, a date and a counterparty.
- 02
Investigate
The agent searches the ledger, the bank feed and the trade system, following what it finds. A near-match by amount leads to a date-shift check. An unknown counterparty leads to a lookup in the master file.
- 03
Conclude
It classifies the break (timing, fee, duplicate, unknown), drafts the journal entry and attaches the evidence it used.
- 04
Review
Timing and fee breaks below a threshold post automatically. Everything else waits for the analyst, who sees the full trace.
Analysts review conclusions instead of gathering evidence. The agent's step limit and tool allowlist keep it inside the three systems it was given.
Stack
What do we typically build it with?
- Models
- Claude or GPT-class models for reasoning steps, smaller models for classification, open-source models on-prem where required
- Harness
- LangGraph or a custom loop with typed tools, step limits and tracing. Model Context Protocol (MCP) servers for tool access where it fits
- Tools
- Read and write connectors for SAP, Salesforce, ServiceNow, databases and internal APIs, each with its own permission scope
- Observability
- Traces and evals in LangSmith or your existing stack (Datadog, OpenTelemetry)
Questions
What people ask about autonomous agents (react)
What is a ReAct agent?
ReAct stands for Reason and Act. The model alternates between reasoning about what to do next and acting through a tool, then reads the result and reasons again. It is the standard pattern for agents that have to work things out rather than follow a script.
How do you stop an autonomous agent from going off the rails?
The harness, not the model, enforces the limits: which tools exist, how many steps are allowed, how much it can spend, and which actions need a person. If the agent hits a limit it stops and hands over with its trace.
Is an autonomous agent slower than a workflow?
Usually, yes. Each loop is a model call plus a tool call, so a run takes seconds rather than milliseconds. That is fine for case work and wrong for real-time paths, which is why we scope the workflow first.
Can it run on our own models?
Yes. The loop is model-agnostic. For regulated environments we deploy open-source models on-prem and run the same evals to check parity.
How do you test an autonomous agent?
We build an evaluation set from real cases during discovery, including the awkward ones, and replay it on every change. The score has to hold before anything ships.
Related
- Read
LangGraph workflow agents
Graph-based workflows where the model decides only at chosen nodes, with checkpointed state and human-in-the-loop interrupts. Example: a resume screening agent.
- Read
RAG agents
Document ingestion with permissions, hybrid keyword and vector retrieval with reranking, and grounded answers with citations or an honest not-found.
- Read
Security and guardrails
How every Unmatrix AI system handles PII and PHI, defends against prompt injection, limits what an agent can do, and is tested before go-live.
Next step
Pick one workflow. We will be at your office in two weeks.
A thirty-minute call to find the right first workflow, followed by a written scoping note. No deck, no pilot.