Architecture

Enterprise AI Agents in Production: What Works Inside Real Systems

Glass cubes on a glowing line pass through a checkpoint gate into a dark cube, with one cube held aside on a branch

An enterprise AI agent that impresses in a demo and one that runs every day inside your ERP use the same model. What separates them is everything around the model: what the agent may read, which actions it may take, who signs off before money moves, how it is stopped, and who notices when its output drifts. Pilots tend to skip those parts because a demo never needs them. Production needs all of them on day one.

What an enterprise AI agent is, in production terms

The joint guidance Principles for the Secure Integration of Artificial Intelligence in Operational Technology, led by CISA and the Australian Signals Directorate in December 2025, defines AI agents as "a type of software that can process data, perform decision-making capabilities, and initiate autonomous actions using AI and ML models." When NIST's Center for AI Standards and Innovation launched its AI Agent Standards Initiative in February 2026, it described agents that "can now work autonomously for hours, write and debug code, manage emails and calendars."

A chatbot that answers a question wrongly produces a bad answer. An agent that acts wrongly posts an invoice, books an appointment or emails a customer. Once an agent can write to a system of record, it needs the same controls you would put around a new employee with system access, and the checks have to run on every single run.

Five checkpoints every production run passes through

Take an accounts payable agent. A supplier emails an invoice for €18,400 against a purchase order approved at €18,000. In production, that one run passes five checkpoints before anything reaches the ledger.

Enterprise AI agents in production
One agent run inside a real system: five checkpoints
Example run
Change the invoice amount, or make the ERP lookup time out.
Example runA supplier emails an invoice for €18,400 against purchase order PO-7731.
€18,400
1 · ReadExtract the fieldsAgentSupplier, amount and PO number read from the PDF.Email text is data, never an instruction.
2 · Look upFetch the contextRead-only toolsPO and supplier record fetched from the ERP, read-only.PO-7731 · approved €18,000
3 · CheckApply the ruleRules engineTolerance is 1% of the PO value. €18,400 is 2.2% over, so the agent cannot post it alone.Needs approval · routed to the AP lead
4 · ActWrite to the ERPTool gatewayAfter approval, one allow-listed action posts the invoice. No other write access.post_invoice · scoped key
5 · RecordLog the runRun logInputs, tool calls, rule, approver and pinned model version.Run #48213 · complete
FailsafeAny step fails or times out: the run stops and the invoice lands in the AP team’s manual queue.
WatchedEvery run feeds the monitors: exception rate, approvals, cost per run, blocked calls.
  1. Read. The agent extracts the supplier, amount and PO number. Anything written in the email or the PDF is treated as data to extract, never as an instruction to follow.
  2. Look up. The agent fetches the purchase order and supplier record through read-only tools. It cannot change what it reads.
  3. Check. A rules engine, not the model, compares the amount with the PO. The tolerance is 1%, the invoice is 2.2% over, so the run routes to the AP lead for approval.
  4. Act. After approval, the agent calls one allow-listed action, post invoice, with a credential scoped to that action alone.
  5. Record. The run log stores the inputs, every tool call, the rule applied, the approver and the model version.

If any step fails or times out, the run stops and the invoice lands in the team's manual queue. Our AI invoice automation service is built on this pattern, with a three-way match and an approver for every exception.

Let the database decide and the model explain

A language model is good at reading messy input and writing a clear sentence. It is unreliable at arithmetic and at applying a threshold the same way twice. NIST AI 800-4, Challenges to the Monitoring of Deployed AI Systems (March 2026), puts it plainly: "AI outputs are typically non-deterministic, meaning the AI may exhibit a range of behaviors under the same input conditions." So in a production agent the numbers and the yes or no come from code, and the model writes around them.

That split is how we built daily MI alerting for IES Limited, a travel insurance group with four brands. Every figure is calculated in PostgreSQL views, n8n decides who needs to know, and the model receives finished numbers and writes the narrative. A verification step names any figure the model was not given. A quote-funnel drop that once took IES 20 working days to notice now reaches the right Teams channel within 30 minutes of the daily pack landing. Replayed against past data, the system caught all 4 earlier incidents that had a baseline to compare against, and IES changed its own alert thresholds 34 times in six weeks without a single developer ticket.

Business rules work the same way. Our post on mapping business rules into the context layer shows how discount limits and approval thresholds move out of the prompt into versioned records with owners.

Connecting agents to your systems with MCP and a context layer

An agent is only as useful as the systems it can reach. The Model Context Protocol, which Anthropic released in November 2024 and contributed to the Linux Foundation's Agentic AI Foundation in December 2025, gives agents one standard way to call tools and fetch data instead of a custom connector per system. n8n, Zapier and most of the platforms in our AI agent platform comparison support it.

The protocol is plumbing. Results depend on what flows through it. A preliminary 2026 NIST SURF student project wrapped a research database in an MCP server and found that "the systemic inclusion of database schemas and handbooks resulted in a statistically significant increase in query success rates." In our builds, describing the data and the rules well has mattered more than which model sits on top. That is the job of an enterprise context layer, and our context layer blueprint covers how to build one.

AI agent security: prompt injection, privileges and tool access

An agent adds a risk a chatbot does not have: it reads content written by people you do not control. NIST AI 600-1, the Generative AI Profile (July 2024), describes it: "Indirect prompt injection attacks occur when adversaries remotely (i.e., without a direct interface) exploit LLM-integrated applications by injecting prompts into data likely to be retrieved." The same document notes that researchers "have already demonstrated how indirect prompt injections can exploit vulnerabilities by stealing proprietary data or running malicious code remotely on a machine."

For the invoice agent, the attack is a PDF carrying hidden text that tells the agent to change the supplier's bank details. No filter catches every phrasing, so the defence is architectural:

  • Least privilege per agent. Each agent gets its own credential, scoped to the tables and actions its job needs. The invoice agent can read POs and post invoices. It cannot edit supplier bank details at all.
  • A tool gateway with an allow-list. Every write passes through one place that checks the action, the parameters and the caller, and blocks anything outside the list.
  • Approval on high-consequence actions. Payments, data exports and changes to master data wait for a named person.
  • A distinct identity in the logs. The CISA-led guidance advises configuring logging "so AI decisions can be tracked for compliance and forensic analysis, and so the logged AI identity is distinct from any typical machine or user identifiers."
  • A pinned model version. A provider's model update can change behaviour without notice, so the version is fixed and changed only after the agent is retested.

If you already run agents on n8n, our n8n consultant service reviews and hardens existing workflows.

Human in the loop and failsafes that actually work

Putting a person on every decision sounds safe and fails in practice. NIST AI 800-4 cites Yampolskiy: "one major issue with human-in-the-loop monitoring is that humans may not be able to keep up with the speed and complexity." Reviewers who approve hundreds of routine items a day stop reading them. Approval belongs where the consequence is high, the rule says so, or the agent's confidence is low, and everything else flows.

The stop button matters as much as the approval queue. The CISA-led guidance tells operators to "establish failsafe mechanisms that enable AI systems to fail gracefully without disrupting critical operations" and to make sure processes can "revert to traditional automation or manual" control. NIST AI 600-1 asks for protocols "to ensure GAI systems are able to be deactivated when necessary." In practice that means a manual queue the team already knows how to work, a switch that pauses the agent without a deployment, and a report for every case the agent declines. In the dental scheduling system we built for a Dutch Odoo implementation partner, covering 200 practitioners and 1,500 appointments a week, any appointment the agent cannot place produces a report explaining why, so a planner picks it up with the reason already in hand.

AI agent monitoring after go-live

Pre-launch testing shows the agent works on the cases you thought of. Monitoring shows what happens on the cases you did not. NIST AI 800-4 groups post-deployment monitoring into six categories: functionality, operational, human factors, security, compliance and large-scale impacts. Each needs its own signals and its own owner.

Enterprise AI agents in production
One invoice agent, week 6 after go-live
Six monitoring categories · NIST AI 800-4
FunctionalityIs it still doing the job?
Fields correct, sampled runs97.8%
Sent to manual queue6.1%
OperationalIs it fast and affordable enough?
Median run time41 s
Model cost per invoice€0.04
Human factorsNeeds actionAre people trusting it the right amount?
Approvals overturned0.4%
Oldest item awaiting approval3 days
SecurityIs anything trying to steer it?
Tool calls blocked by gateway2
Instruction-like text in inputs5 flagged
ComplianceCan we prove what it did?
Runs with a complete log100%
Reads outside its scope0
Large-scale impactsWhat is it changing downstream?
Supplier payment disputesFlat
AP team hours on intakeDown
Every technical number is healthy. The problem is in human factors: an approval has waited three days.
Illustrative figures. The categories come from NIST AI 800-4 (March 2026); the signals under each are the ones we would set for an invoice agent.

In the example board, every technical number is healthy and the real problem sits in human factors: an approval has waited three days, so the agent is fast and the process around it is slow. That is the kind of issue a dashboard of model metrics never shows.

NIST is candid about what is still unsolved. The report describes a possible "monitorability tax", paying for slightly less capable models or more expensive inference to keep reasoning observable. It names a privacy versus granularity trade-off, because detailed logs capture personal data. And it notes that the "appropriate metrics to capture is not standardized." Until standards catch up, each agent needs monitors chosen for its own job. Our AI managed services cover that run loop after launch, and enterprise AI governance covers the policy and audit controls around it.

How to move an AI agent from pilot to production

  1. Pick one workflow with a clear owner and a measurable cost. Our automation ROI calculator and AI use case canvas help rank candidates.
  2. Map the rules before you write a prompt. Decide which checks belong in code and which approvals need a person.
  3. Build on your real data. A staging copy of your ERP or CRM exposes the malformed records and edge cases a demo dataset hides. Our enterprise AI consulting prototypes run this way from week one.
  4. Set the monitors and the failsafe before launch. Agree the six categories, the thresholds and who is alerted, and test the manual fallback.
  5. Price it at production volume. Model cost per run, multiplied by real volume, is easy to underestimate. The AI agent cost estimator gives a first number.
  6. Run it, review it, improve it. A weekly review of exceptions and overrides shows which rules, prompts or thresholds need to change.

If you are not sure where to start, the AI readiness scorecard takes a few minutes, and our AI Audit ends with a 90-day roadmap of the workflows worth automating first. How we work shows who owns each stage.

Frequently Asked Questions About Enterprise AI Agents

What is an enterprise AI agent?

An enterprise AI agent is software that uses a language model to read inputs, decide on a next step and take actions in business systems such as an ERP, CRM or inbox. What makes it enterprise grade is the layer around the model: scoped access, rules that run in code, approvals for high-consequence actions, logging and monitoring.

How do you deploy AI agents in production safely?

Give each agent least-privilege credentials, route every write through an allow-listed tool gateway, keep thresholds and calculations in a rules engine or database, require approval for payments and data changes, pin the model version, log every run and build a manual fallback before launch.

What is indirect prompt injection?

NIST AI 600-1 describes it as an attack where adversaries inject prompts "into data likely to be retrieved", such as an email, a document or a web page the agent reads. The agent may then follow the hidden instructions. The defence is to treat retrieved content as data and limit what the agent is able to do, so a successful injection has nothing harmful to trigger.

What should you monitor on an AI agent?

NIST AI 800-4 groups monitoring into functionality, operational, human factors, security, compliance and large-scale impacts. For a working agent that means sampled accuracy, exception rate, run time and cost, override rate and approval wait times, blocked tool calls, log completeness and the downstream effect on the team.

Do AI agents need a human in the loop?

For high-consequence actions, yes. Reviewing every routine decision tends to fail, because reviewers cannot keep up and start approving without reading. Route the decisions that carry risk to a named approver and let the rest run under monitoring.

Put your first agent into production

In a 30-minute discovery call you speak directly with an AI engineer about the workflow you want to automate, the systems it touches and the controls it needs. If an audit is the right first step, we will tell you and scope it on the call.

Book a Discovery Call

Sources

  1. CISA and the Australian Signals Directorate's ACSC, with NSA, FBI and the cyber agencies of Canada, Germany, the Netherlands, New Zealand and the UK (December 2025). Principles for the Secure Integration of Artificial Intelligence in Operational Technology.
  2. NIST Center for AI Standards and Innovation (February 2026). Announcing the AI Agent Standards Initiative for Interoperable and Secure Innovation.
  3. NIST (March 2026). AI 800-4, Challenges to the Monitoring of Deployed AI Systems.
  4. NIST (July 2024). AI 600-1, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile.
  5. NIST Summer Undergraduate Research Fellowship (August 2026). 2026 SURF Colloquium abstract book. Preliminary student research.
  6. Anthropic (November 2024). Introducing the Model Context Protocol. Linux Foundation (December 2025). Agentic AI Foundation announcement.

Ready to get started?

A 30-minute discovery call. You bring the process; we bring the plan.

Book a Discovery Call