
An enterprise AI agent that impresses in a demo and one that runs every day inside your ERP use the same model. What separates them is everything around the model: what the agent may read, which actions it may take, who signs off before money moves, how it is stopped, and who notices when its output drifts. Pilots tend to skip those parts because a demo never needs them. Production needs all of them on day one.
The joint guidance Principles for the Secure Integration of Artificial Intelligence in Operational Technology, led by CISA and the Australian Signals Directorate in December 2025, defines AI agents as "a type of software that can process data, perform decision-making capabilities, and initiate autonomous actions using AI and ML models." When NIST's Center for AI Standards and Innovation launched its AI Agent Standards Initiative in February 2026, it described agents that "can now work autonomously for hours, write and debug code, manage emails and calendars."
A chatbot that answers a question wrongly produces a bad answer. An agent that acts wrongly posts an invoice, books an appointment or emails a customer. Once an agent can write to a system of record, it needs the same controls you would put around a new employee with system access, and the checks have to run on every single run.
Take an accounts payable agent. A supplier emails an invoice for €18,400 against a purchase order approved at €18,000. In production, that one run passes five checkpoints before anything reaches the ledger.
If any step fails or times out, the run stops and the invoice lands in the team's manual queue. Our AI invoice automation service is built on this pattern, with a three-way match and an approver for every exception.
A language model is good at reading messy input and writing a clear sentence. It is unreliable at arithmetic and at applying a threshold the same way twice. NIST AI 800-4, Challenges to the Monitoring of Deployed AI Systems (March 2026), puts it plainly: "AI outputs are typically non-deterministic, meaning the AI may exhibit a range of behaviors under the same input conditions." So in a production agent the numbers and the yes or no come from code, and the model writes around them.
That split is how we built daily MI alerting for IES Limited, a travel insurance group with four brands. Every figure is calculated in PostgreSQL views, n8n decides who needs to know, and the model receives finished numbers and writes the narrative. A verification step names any figure the model was not given. A quote-funnel drop that once took IES 20 working days to notice now reaches the right Teams channel within 30 minutes of the daily pack landing. Replayed against past data, the system caught all 4 earlier incidents that had a baseline to compare against, and IES changed its own alert thresholds 34 times in six weeks without a single developer ticket.
Business rules work the same way. Our post on mapping business rules into the context layer shows how discount limits and approval thresholds move out of the prompt into versioned records with owners.
An agent is only as useful as the systems it can reach. The Model Context Protocol, which Anthropic released in November 2024 and contributed to the Linux Foundation's Agentic AI Foundation in December 2025, gives agents one standard way to call tools and fetch data instead of a custom connector per system. n8n, Zapier and most of the platforms in our AI agent platform comparison support it.
The protocol is plumbing. Results depend on what flows through it. A preliminary 2026 NIST SURF student project wrapped a research database in an MCP server and found that "the systemic inclusion of database schemas and handbooks resulted in a statistically significant increase in query success rates." In our builds, describing the data and the rules well has mattered more than which model sits on top. That is the job of an enterprise context layer, and our context layer blueprint covers how to build one.
An agent adds a risk a chatbot does not have: it reads content written by people you do not control. NIST AI 600-1, the Generative AI Profile (July 2024), describes it: "Indirect prompt injection attacks occur when adversaries remotely (i.e., without a direct interface) exploit LLM-integrated applications by injecting prompts into data likely to be retrieved." The same document notes that researchers "have already demonstrated how indirect prompt injections can exploit vulnerabilities by stealing proprietary data or running malicious code remotely on a machine."
For the invoice agent, the attack is a PDF carrying hidden text that tells the agent to change the supplier's bank details. No filter catches every phrasing, so the defence is architectural:
If you already run agents on n8n, our n8n consultant service reviews and hardens existing workflows.
Putting a person on every decision sounds safe and fails in practice. NIST AI 800-4 cites Yampolskiy: "one major issue with human-in-the-loop monitoring is that humans may not be able to keep up with the speed and complexity." Reviewers who approve hundreds of routine items a day stop reading them. Approval belongs where the consequence is high, the rule says so, or the agent's confidence is low, and everything else flows.
The stop button matters as much as the approval queue. The CISA-led guidance tells operators to "establish failsafe mechanisms that enable AI systems to fail gracefully without disrupting critical operations" and to make sure processes can "revert to traditional automation or manual" control. NIST AI 600-1 asks for protocols "to ensure GAI systems are able to be deactivated when necessary." In practice that means a manual queue the team already knows how to work, a switch that pauses the agent without a deployment, and a report for every case the agent declines. In the dental scheduling system we built for a Dutch Odoo implementation partner, covering 200 practitioners and 1,500 appointments a week, any appointment the agent cannot place produces a report explaining why, so a planner picks it up with the reason already in hand.
Pre-launch testing shows the agent works on the cases you thought of. Monitoring shows what happens on the cases you did not. NIST AI 800-4 groups post-deployment monitoring into six categories: functionality, operational, human factors, security, compliance and large-scale impacts. Each needs its own signals and its own owner.
In the example board, every technical number is healthy and the real problem sits in human factors: an approval has waited three days, so the agent is fast and the process around it is slow. That is the kind of issue a dashboard of model metrics never shows.
NIST is candid about what is still unsolved. The report describes a possible "monitorability tax", paying for slightly less capable models or more expensive inference to keep reasoning observable. It names a privacy versus granularity trade-off, because detailed logs capture personal data. And it notes that the "appropriate metrics to capture is not standardized." Until standards catch up, each agent needs monitors chosen for its own job. Our AI managed services cover that run loop after launch, and enterprise AI governance covers the policy and audit controls around it.
If you are not sure where to start, the AI readiness scorecard takes a few minutes, and our AI Audit ends with a 90-day roadmap of the workflows worth automating first. How we work shows who owns each stage.
An enterprise AI agent is software that uses a language model to read inputs, decide on a next step and take actions in business systems such as an ERP, CRM or inbox. What makes it enterprise grade is the layer around the model: scoped access, rules that run in code, approvals for high-consequence actions, logging and monitoring.
Give each agent least-privilege credentials, route every write through an allow-listed tool gateway, keep thresholds and calculations in a rules engine or database, require approval for payments and data changes, pin the model version, log every run and build a manual fallback before launch.
NIST AI 600-1 describes it as an attack where adversaries inject prompts "into data likely to be retrieved", such as an email, a document or a web page the agent reads. The agent may then follow the hidden instructions. The defence is to treat retrieved content as data and limit what the agent is able to do, so a successful injection has nothing harmful to trigger.
NIST AI 800-4 groups monitoring into functionality, operational, human factors, security, compliance and large-scale impacts. For a working agent that means sampled accuracy, exception rate, run time and cost, override rate and approval wait times, blocked tool calls, log completeness and the downstream effect on the team.
For high-consequence actions, yes. Reviewing every routine decision tends to fail, because reviewers cannot keep up and start approving without reading. Route the decisions that carry risk to a named approver and let the rest run under monitoring.
In a 30-minute discovery call you speak directly with an AI engineer about the workflow you want to automate, the systems it touches and the controls it needs. If an audit is the right first step, we will tell you and scope it on the call.
A 30-minute discovery call. You bring the process; we bring the plan.
Book a Discovery Call