Travel tips and city guides

Taking AI Agents to Production: Middleware, Durable Workflows and Security

AI and Agents ·

Taking AI agents to production — abstract illustration of multimodal AI | Aksiyon Soft

Why do AI agents that shine in demos break in production?

AI agents look flawless in a demo: you define a few tools, give the model a goal, and the agent reads an email, finds the customer in the CRM and drafts a quote. The trouble starts with real users, real data and real permissions. When the process crashes halfway, the agent forgets where it was, it may issue the same invoice twice, it can be talked into an unauthorized action by instructions hidden inside a PDF, and at month end nobody can explain why the token bill tripled.

This guide is for IT leaders, software architects and product owners who want to connect agents to ERP, CRM, email and internal portals. The topic is not how smart the model is but the middleware chain around the agent loop: identity and authorization, policy and guardrails, budget limits, a durable step store, human approval, an audit log and observability. The announcements published in the first week of October 2026 by Anthropic, Kong, MongoDB and the open-source community all point at this exact layer; below we map which product fills which gap and what remains your responsibility.

In short

  • An AI agent in production is not model + tools; it is model + tools + a middleware chain.
  • Without durable execution (checkpoints), long-running workflows that wait for human approval are not reliable.
  • Prompt injection and excessive agency (OWASP LLM01 and LLM06) are the two most common risks; least privilege and approval steps are the answer.
  • MCP connects agents to tools and A2A connects agents to each other; neither replaces controlled access to ERP and CRM.
  • Writes to systems of record should always go through an idempotent integration layer.

• • •

What separates a demo agent from a production agent?

In a demo the agent runs in one session, for one user, usually with admin rights. In production the same agent has to act on behalf of hundreds of users, with different roles, in workflows that last hours or even days. The mistake we see most often in projects is shipping demo code “with a bit of hardening”; what is missing is not a few try/catch blocks but an entire runtime layer.

Lost state

If the agent loop lives in memory, everything is gone when the process restarts. A purchase request awaiting approval, a half-finished data transfer or a reply waiting to be sent all have to start over, which creates the risk of duplicate records.

Permissions and trust boundaries

The broad API key handed to a demo agent can become a disaster in production. Agents need per-user permissions, separate read and write tools, and a human sign-off on critical actions.

Cost and blind spots

The agent loop keeps iterating on its own; without limits, a single faulty task can trigger hundreds of model calls. If you cannot see which user, in which task, called which tool how many times, you can manage neither cost nor failures.

Source code on a dark screen — coding middleware and tool permissions for AI agents
Moving demo code to production is not about adding error handlers; it means building a separate runtime layer for state, permissions and cost.

• • •

What is an agent middleware chain?

The middleware pattern we know from web apps maps directly onto agents: every request or event passes through a chain of layers, each layer does one job and hands over with next(). The difference is that in agent land a “request” is not only an HTTP call; a model call, a tool call, a permission request and the end of a turn are all events.

Anthropic’s mods, shipped on October 1, 2026 in Claude Code 2.1.287, show how mainstream the idea has become. According to the official announcement, a mod can rewrite a prompt before it reaches the model, block, rewrite or retry a tool call, approve or deny permission requests, and redact secrets from tool output before Claude reads it. When several mods hook the same event they run in load order: the first mod sees the event first and the result last. Anthropic’s own warning matters too: mods are not sandboxed and run with the same access to your machine as Claude Code itself.

Diagram: AI agent middleware chain — observability, identity, guardrails, budget, human approval, durable step store and tool router
The agent middleware chain: requests travel from the outer layers inwards and responses return the same way; only the integration layer touches enterprise systems.

The seven layers of the chain

The order we use in our own projects, from the outside in, looks like this. Observability sits on the outside so that every other layer’s decisions end up in the trace. Next comes identity and authorization: the agent always acts on behalf of a user or a service account. The policy and guardrail layer screens inputs and tool output for prompt injection and data leakage.

The budget and rate-limit layer caps turns, tokens and tool calls. The human approval layer pauses the flow on risky actions. The durable step store persists each step result so the run can resume after a crash. At the core sits the tool router, which decides which tool is served by an MCP server and which task is delegated to another agent over A2A. The audit log lives next to the chain as a separate, append-only store.

• • •

How do you build durable workflows with human approval?

Most enterprise agent tasks take hours, not minutes: a purchase request waits for a manager, a return waits for a warehouse check. Nobody can guarantee the server will not restart in the meantime. The fix is to build the agent loop on durable execution: each step result is written to persistent storage, and after a restart completed steps are read from the log instead of being executed again.

LangGraph: checkpointers and interrupts

The LangGraph documentation covers this with two pieces. A checkpointer saves graph state at every step, and human-in-the-loop, time travel and fault tolerance are built on top of it. Calling interrupt() pauses execution, saves state and waits indefinitely until you resume. The docs define three durability modes — exit, async and sync — and sync, which writes every checkpoint before the next step starts, gives the strongest guarantee. Note that the in-memory checkpointer loses everything on restart; for production the docs point to a PostgreSQL-backed checkpointer.

The LangChain team walks through how interrupt-based approval relies on the checkpointer and how a paused agent run is resumed.

Kong Volcano and Obelisk: the same idea as a platform

Kong’s Volcano, announced on September 30, 2026, brings durable compute, branchable PostgreSQL databases, file storage, user authentication and real-time services under one roof. Per the Volcano docs, a durable function records each step with ctx.step, can keep working in the background for up to 366 days, and can wait for a person’s approval with ctx.waitUntil. On the open-source side, Obelisk 0.42 (October 4, 2026) suggests running the agent itself as a durable workflow, with model calls and tools as activities, which the release notes call useful for enterprise agents that wait on people or external systems.

Diagram: durable agent workflow — read from ERP, LLM draft, human approval, idempotent ERP write and checkpoint store
Each step result lands in the checkpoint store; after a crash, completed steps replay from the log and the ERP write is protected by an idempotency key.

• • •

How do MCP and A2A connect agents to tools and to each other?

The Model Context Protocol (MCP) standardizes how an agent connects to tools and data sources. A2A (Agent2Agent) is an open protocol hosted by the Linux Foundation that lets independent agents discover each other, delegate work and exchange results. The A2A documentation describes the two as complementary: MCP covers agent-to-tool and A2A covers agent-to-agent communication.

MongoDB shows how fast enterprises are adopting these standards. On September 29, 2026 the company launched Atlas Agent Engine in public preview, positioning it as a unified execution, memory and governance layer for production agents. The product page says it uses MCP for tools, A2A for agent-to-agent delegation and OpenTelemetry for traces. Before wiring a preview product into critical processes, we recommend pinning down SLA and pricing terms separately.

Protocols make connections easier, but they do not solve authorization. If an MCP server grants write access to every table in your ERP, the problem is the design, not the protocol. That is why MCP tools should talk to an integration middleware layer that enforces permissions, never to the database directly.

• • •

How do you secure AI agents against prompt injection?

LLM01 Prompt Injection tops OWASP’s 2025 Top 10 for LLM applications. The risk is not limited to what users type; it also hides in the emails, web pages, documents and tool responses the agent reads. LLM06 Excessive Agency on the same list stresses that agentic architectures make the issue more critical: excessive functionality, permissions and autonomy turn a manipulated output into a damaging action.

What to do in practice

OWASP’s guidance and our field experience meet at the same point: minimize tools and their permissions, avoid open-ended tools such as free-form SQL or shell commands, run the agent with the user’s own permissions, require human approval for high-impact actions, and enforce authorization in the downstream system rather than trusting the agent. Our article on RBAC access model design is a good starting point for the role model.

Team at a whiteboard covered in sticky notes — designing human approval and permission rules for AI agents
Which actions need approval is a business decision, not a technical one; thresholds and approver roles belong in the discovery workshop.

• • •

Which product covers which layer?

The table below maps the notable announcements from early October 2026 onto the layers of the agent middleware chain. None of them alone means “production-ready AI agents”; each one strengthens one or more links of the chain.

Product / approachStatusMain layer it addressesWhat to watch
Claude Code mods (2.1.287)Shipped October 1, 2026Middleware chain for agent-loop events: prompts, tool calls, permissionsMods are not sandboxed; install only mods from sources you trust
Kong VolcanoAnnounced September 30, 2026Durable compute, branchable PostgreSQL with vector support, authenticationSandboxed compute follows after the launch; review data location separately
MongoDB Atlas Agent EngineSeptember 29, 2026, public previewExecution, memory and governance; MCP, A2A and OpenTelemetryPreview status; confirm SLA and pricing
LangGraphOpen-source frameworkDurable execution via checkpointers, human approval via interruptsDurability mode and persistent checkpointer are your choice
Obelisk 0.42October 4, 2026Deterministic durable workflow runtime; agents as workflowsRelease includes breaking configuration and JavaScript API changes
Your own integration layerNeeded in every projectAuthorized ERP/CRM writes, idempotency, outbox, audit logStays with you regardless of vendor choice

• • •

How should AI agents connect to ERP and CRM?

An agent should never connect to systems of record directly. The right model is an agent that calls a limited set of documented tools, which in turn reach ERP, CRM and email through an integration layer. That layer applies authorization, data validation, rate limits and error handling centrally.

Two patterns are non-negotiable for writes. The first is an idempotency key: when the durable workflow replays a step, the same order is not created twice. The second is the outbox pattern: the database change and the message to the external system are stored in the same transaction. Our article on the outbox pattern, idempotency and retries covers the details. If you want to plug Türkiye-hosted models into the same layer, our guide on connecting EVREN API to enterprise software gives a hands-on example.

When you use more than one model, connecting the agent to an LLM router instead of a single model pays off in budget control, data-sensitivity rules and redundancy. We cover that layer in the AI Ops layer, LLM routers and KVKK. To feed agent traces, cost and error rates into your existing monitoring, see observability in enterprise applications.

• • •

Production-readiness checklist

Before AI agents go live, you should have a written answer for every item below. Any item answered with “we will look at it later” will surface in the first serious incident.

  • Which user or service account does the agent act for, and are its permissions minimized?
  • Are read and write tools separated, and are open-ended tools such as free-form SQL or shell commands removed?
  • Is there a per-task budget for turns, tokens and tool calls?
  • Are the actions that require human approval, the approver roles and the timeout rules written down?
  • Is the durable step store a persistent database, and has a crash-and-restart scenario been tested?
  • Is every step that writes to ERP or CRM protected by an idempotency key?
  • Are tool outputs and external documents screened for prompt injection, and are secrets masked?
  • Is every model and tool call traced, and is the audit log kept in tamper-resistant storage?
  • Is there a fallback model or a safe-stop behavior when the model provider goes down?
Anthropic’s team discusses when an agent truly adds value, when a simple workflow is enough, and why results have to be measured.

• • •

How can Aksiyon Soft help?

Through our API and integration service we design the integration layer, permission model and audit log that connect agents to ERP, CRM and internal systems. When your processes do not fit off-the-shelf products, we build the agent middleware chain on your own infrastructure through custom software development. You can review our architecture approach on the API and data integration platform page.

Our delivery model starts with discovery: which tasks suit an agent, which actions need approval and which systems will be touched are put in writing. We then build an MVP for one process, demo the working flow at the end of every sprint, and follow go-live with a hypercare period and SLA-backed support. We are headquartered in Samsun and work remotely with teams across Türkiye, with planned on-site visits when the project needs them.

Frequently asked questions

How are AI agents different from classic automation?

Classic automation follows predefined steps; an agent decides which tool to call and in what order based on a goal. That flexibility also opens the door to unpredictable behavior when limits and approval steps are missing.

Does every agent need a durable workflow engine?

Not for tasks that finish in seconds and have no side effects. Any task that runs long, waits for human approval or writes to enterprise systems needs checkpoint-based durable execution.

Does using an MCP server solve security?

No. MCP standardizes how tools are connected; which tool is exposed to whom, with which permissions, is your design decision. Authorization must be enforced in the target system and the integration layer.

Can prompt injection be fully prevented?

With today’s models it does not look fully preventable. The goal is layered risk reduction: input and output screening, least privilege, human approval on critical actions and separate authorization in the target system.

How long does it take to put AI agents into production?

An MVP that starts with one process and a small tool set can go live in a few sprints. Programs with ERP, CRM and multi-step approvals can take months depending on integration scope; we share a written timeline after discovery.

What happens if the underlying model changes?

If the agent talks to an OpenAI-compatible LLM router rather than a single model, switching models is mostly a configuration change. You still need to rerun your evaluation tests against the new model.

Let's talk about your project

If your agent idea works in a demo but the path to production is unclear, send us a short summary via the contact page. We will work out together which process suits an agent, which middleware layers you need and what the integration scope looks like.

Sources

Related posts

Directions