E
very agent demo looks impressive. The model reasons through a problem, calls a tool, checks the result, adjusts its approach, gets the answer. Clean. Autonomous. Feels like science fiction shipping as a product.Then you try to run it in production. The agent calls the wrong tool. It loops. It confidently completes a task it didn't actually complete. It works perfectly on the demo dataset and fails on real user inputs in ways that are hard to reproduce and harder to debug.
AI agents production architecture is a genuinely different problem from building a standard LLM feature. The decision to commit to an agentic approach is one of the more consequential architecture choices a CTO can make right now mostly because the cost of getting it wrong is high and the signals during development are misleading. Here's what I'd want to understand before making that call.
What "agent" actually means in a production context
The word is overloaded. When I say agent in a production context, I mean a system where a model makes sequential decisions choosing which action to take next based on previous results with access to tools it can call autonomously, operating toward a goal without human confirmation at each step.
That last part is what makes agents different from a standard LLM call with function calling. In a non-agentic setup, the model suggests an action, your code confirms it, executes it, returns the result. Control stays with your code. In an agentic setup, the model drives the loop. It decides what to do next. Your code executes what it asks.
That shift in control is where most production problems come from.
The practical spectrum looks like this. On one end: a model that picks from a fixed menu of actions and calls exactly one per request. That's barely an agent it's structured output with a function call. On the other end: a model running an open-ended research loop, spawning sub-agents, writing and executing code, sending emails. That's a genuine agent and it comes with genuine production risk.
Most products that claim to need the second thing actually work fine with something much closer to the first. Worth establishing where on that spectrum your use case actually sits before building anything.
The four failure modes that kill agentic systems in production
These aren't edge cases. They're common enough that if you ship an agent without planning for them, you'll hit at least two.
Loops. The agent enters a state where its next action keeps producing the same result, which it interprets as a reason to try the same action again. Without explicit loop detection a step counter, a state hash, a maximum iteration limit this runs until you hit a timeout or exhaust your token budget. I've seen this produce $200 of API spend on a single malformed user request.
Confident incompletion. The agent reports success on a task it didn't complete. It called the right tool. The tool returned an ambiguous response. The model interpreted ambiguity as success and reported done. This is the most dangerous failure mode for any agent touching external systems databases, APIs, payment flows because the audit trail looks clean and the failure is silent.
Tool misuse. The agent calls a tool with parameters that are technically valid but semantically wrong. A search tool called with an overly broad query that returns noise, which the agent then reasons on top of incorrectly. A database write called with a field value that passes schema validation but makes no business sense. Regular software has types and constraints to catch this. Agents have a model that's confident it knows what it's doing.
Compounding errors. In a multi-step agent loop, an error in step two poisons everything that follows. The agent doesn't backtrack. It reasons forward from a corrupted state and produces a final result that's wrong in ways that are hard to trace back to the original mistake. The longer the chain, the worse this gets.
The engineering response to all four is the same in structure: bounded execution, explicit checkpoints, and deterministic validation at tool boundaries. But the specific implementation depends on your use case, which is why "we'll handle edge cases after launch" is the wrong approach for agentic systems.
The architecture decisions that actually matter
Assuming you've decided an agentic approach is right for your use case, three architectural decisions determine most of your production reliability.
How much autonomy does the agent have? Every tool the agent can call is a potential blast radius. An agent that can read data is low risk. An agent that can write data is higher risk. An agent that can send communications, trigger payments, or modify production records is highest risk. The principle: give the agent the minimum tool access required to accomplish the task. Add tools incrementally as trust is established through observation in production, not during development.
Where are the human checkpoints? For any action that's irreversible or has meaningful external consequences, the agent should pause and confirm before executing. This feels like it defeats the purpose of automation. It doesn't. An agent that handles 90% of a workflow autonomously and pauses on the 10% that's genuinely ambiguous or high-stakes is far more useful in production than an agent that tries to handle 100% and occasionally does something catastrophic. [→ API design for high-throughput systems] has relevant thinking on idempotency the same principles apply when designing agent tool calls.
How do you observe what the agent is actually doing? Standard request/response logging doesn't work for agents because the interesting information is in the reasoning trace, not just the inputs and outputs. You need to log every step of the agent loop: which tool was called, with what parameters, what it returned, and what the model decided to do next. Without this, debugging a production failure in an agentic system is archaeology you're trying to reconstruct what happened from whatever artifacts the agent left behind.
The OpenTelemetry tracing model maps reasonably well to agent steps if you treat each tool call as a span. Several LLM observability tools LangSmith, Langfuse, Arize are purpose-built for this. The choice of tooling matters less than the habit of instrumenting from day one.
The orchestration question: build or use a framework?
LangChain, LlamaIndex, CrewAI, AutoGen the frameworks are everywhere and the ecosystem moves fast. The question for a production team isn't which framework is best. It's whether you need one at all.
Frameworks are good for: getting to a working prototype fast, handling boilerplate around tool calling and memory management, and providing abstractions that make complex agent topologies (parallel agents, supervisor agents, handoffs) easier to express in code.
Frameworks are bad for: production debugging, because the abstraction layers hide what's actually happening; performance tuning, because the defaults are optimised for flexibility not efficiency; and stability, because the ecosystem moves fast enough that upgrading dependencies becomes a risk.
My general recommendation: use a framework to validate that the agentic approach works for your use case. Once it does, evaluate whether you want to carry that framework dependency into production or reimplement the specific orchestration your use case needs in vanilla code. The reimplement is almost always simpler than it looks, because production agents tend to have much more constrained behaviour than the demos suggest.
Real-world example: document processing agent at a Jakarta law firm SaaS
A legal tech startup in Jakarta had built a contract review feature. Initial version: upload a contract, model returns a structured summary. Worked. Users wanted more specific clause extraction, cross-referencing against a library of standard terms, flagging deviations.
The team built an agent. It would parse the document, identify clause types, query the standard terms library for each one, compare, and produce a deviation report. In development, impressive. In production, three problems within the first two weeks.
First: the clause identification step occasionally mis-categorised a clause, which caused the subsequent library query to retrieve irrelevant terms, which caused the comparison to flag non-existent deviations. The agent completed successfully. The report was wrong. Users didn't know until a lawyer reviewed it manually.
Second: long contracts 80+ pages caused the agent to exceed its context window mid-loop. The agent would fail partway through without a clear error, returning a partial report that looked complete.
Third: on degraded API days, intermediate steps would time out, the agent would retry from the beginning, and the user would receive duplicate reports.
The fixes were structural. Clause identification became a validated step with a confidence threshold below the threshold, the clause gets flagged for human review rather than being passed to the next step. Long contracts get chunked before the agent loop starts, with the agent operating on chunks and a separate synthesis step combining results. Each agent run gets a unique idempotency key so retries produce a single report regardless of how many times the loop runs.
None of these are AI problems. They're distributed systems problems that happen to involve an LLM in the loop.
FAQ
Q: How do I know if my use case actually needs an agent, or if a simpler LLM call would work?
A: If you can write out the steps your feature takes as a fixed sequence always step one, then step two, then step three you probably don't need an agent. Agents add value when the next step genuinely depends on the result of the previous step in a way you can't anticipate at design time. If the branching logic is finite and knowable, implement it explicitly in code and call the model for the intelligence parts. You'll have a more reliable system.
Q: What's a reasonable scope for a first production agent?
A: Narrow the tool access to read-only operations first. No writes, no external communications, no irreversible actions. Get the core reasoning loop working reliably on your real production inputs. Add write capabilities incrementally with explicit confirmation checkpoints. Most teams that have production incidents with agents skipped the read-only phase.
Q: How do we handle the case where the agent produces a wrong result confidently?
A: This is the hardest problem and there's no complete solution. The partial solutions: output validation against a schema or set of constraints that catches structurally invalid results; sampling run the agent multiple times on high-stakes inputs and compare outputs; human review on a percentage of results, especially in early production; and observable reasoning traces so that when a wrong result is reported, you can reconstruct what happened.
Q: Multi-agent systems do they actually work in production?
A: Sometimes. The pattern a supervisor agent that routes tasks to specialised sub-agents can work well when each sub-agent has a genuinely narrow and well-defined task. The failure mode is complexity: each agent-to-agent handoff is a potential point of information loss, and debugging a failure across three agents is significantly harder than debugging a single agent. Start with one agent that does more before splitting into multiple agents that do less.
Q: What should I ask my engineering team before we commit to building an agent?
A: Four questions. What happens when the agent loops? What happens when a tool call fails mid-run? What's the maximum number of steps the agent can take before we cut it off? And where can we see a log of exactly what the agent did on any given run? If they don't have answers to all four, the architecture isn't ready for production yet.
The gap between an agent that works in a demo and one that works reliably in production is mostly an engineering discipline gap, not a model capability gap. The model is good enough. The systems around it the observability, the guardrails, the fallbacks, the checkpoints are what most teams underinvest in. If you're making the decision now about whether to commit to an agentic architecture, read the [→ AI-Native Startup Architecture guide] first it covers the broader infrastructure context this decision sits inside.
External Documentation:
- [LangSmith Docs] LLM observability and tracing platform built for agentic workflows; useful for production agent debugging.
- [Langfuse Docs] Open-source LLM observability with agent trace support; self-hostable for data-sensitive environments.