AI Agents Failures: Architecture at Fault
Most AI Agents Fail in Production Because They Are Built Backwards
Good models do not save a bad architecture, and most teams learn this the hard way. When an AI agent system fails in production, it isn't dramatic. There is no crash or error message. The system simply keeps running and producing results that seem reasonable until someone reads them carefully and notices there is a problem.
When a team decides to look into the situation, it can take two days of debugging to understand what is happening. Surprisingly, the model isn't hallucinating, and the input-output tools are producing the correct results. The problem, once it is finally found, is architectural. The model and tools are correctly configured, but the idea is that reasoning would tie everything together, which, as you can guess, obviously fails. It turns out that reasoning doesn't do that kind of thing.
This scenario explains why so many AI agents that work in demonstrations don't really survive real-world use. It's not a capacity issue. It's an architectural issue. The pattern is a familiar one: systems built top-down, from the goal to the tools to the model, with the silent assumption that intelligent behavior will fill in the gaps. This assumption is what "built backwards" means. And it's more common than most teams realize until something breaks.
Agents Are Not Entities. They Are Systems.
An AI agent in production is not a single intelligent thing. It is rather a set of interacting pieces with different responsibilities, failure modes, and levels of observability. The LLM (large language model) is one of those components, not the whole system. Just a piece of it.
This may seem obvious when stated aloud. But the framework of the "autonomous agent" that dominated 2023 and most of 2024 constantly pushed engineers toward a different mental model: an entity, a reasoning loop, all managed by the model. All you need are tools, a good system prompt, and the hope that everything will fit together correctly.
In contrast, engineers who have shipped real AI-based products rarely describe their systems this way. What they describe looks much more like a distributed systems architecture. Not because they read a book on design patterns, but because they have been burned enough times to start structuring their workflow more seriously.
Building top-down, starting with "what should this agent do" and working backward to the tools and prompts, is quick to start. It's also how you end up with a system where the model is responsible for too many things, and where nothing is individually debuggable. The architecture was decided by the goal, not by engineering requirements. That's the "backwards" part.
So, What Does It Really Take for a Production System?
The abstract version is easy to accept. Here's what it actually looks like. Every AI system in production that works correctly has something like a decision layer, whether the team named it that or not. This is the part where the model lives and does its real work.
The instinct is to push everything into this layer: analyzing requests, managing memory, handling retries, resolving tool failures. This is acceptable if you are working in a Jupyter notebook. In production, under load, with real users, it becomes the part of the system where everything is everyone's fault, and most of the time, nothing can be debugged.
The decision layer should do just one thing well: decide what to do next, given a certain context already prepared for it. That's the whole job. Who prepares the context? Something else. Who acts on the decision? Also something else.
That "something else" is the orchestration layer, and in most well-built systems, it's simply code: conditionals, asynchronous executors, retry management, queue routing, maybe even a state machine depending on the complexity of the workflow. Instead of expecting the model to do everything, it is treated as another component. Here, standard code does the heavy lifting with state and tools, so the LLM only has to worry about making the next decision.
Many teams turn to frameworks here because bare orchestration code seems too simple, as if there should be more infrastructure. There usually isn't. The less magic this layer contains, the more quickly bugs are found when they appear. And they will appear.
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
One example: on a project where orchestration lived inside the execution model of a framework, something was retrying tool calls in a way that corrupted downstream state. It took two days to find the problem. Two days for a bug that could have been resolved in no time if the retry logic had been three lines of Python written in-house.
This leads to the tools and execution layer, where all the communication happens. This is where things talk to the outside world. This layer generally has one job, which is to take a well-defined input and produce a predictable output.
But the failure that keeps coming back is tools trying to be helpful by doing more than one thing. A single function that calls an API, updates a cache, and does other things. In a setup like that, when it breaks, you don't know where. Even when you try to replace the API, you're untangling logic that shouldn't have been intertwined in the first place.
Memory and state are the areas that deserve the hardest push, since that's where most teams are the least prepared. Most teams think of memory as "what the model knows." The most important question is what the system knows, and whether that knowledge is up to date.
It can take an afternoon to debug what seems to be a simple "model hallucination." The model keeps referring to user preferences, which were, however, updated twenty minutes earlier. This isn't a model problem. It's a system problem. And it's surprisingly common.
In multi-agent systems, in particular, shared state is where subtle failures develop. One agent updates something. The others don't know. Everyone confidently moves forward in slightly different directions. The output seems almost correct, which is almost worse than seeming incorrect.
And then there's evaluation and observability, which almost everyone always puts off until something goes wrong. The difference to keep in mind is that logging tells you what happened. Observability tells you whether what happened was correct. In a deterministic system, these two things are almost identical.
In an AI system, that's not the case. You need to be able to trace the specific request from start to finish, including what information the model had to consider, what decision it made, what external API call it invoked, and how it acted on its response.
Building the Right Way
It starts with the top-down approach: you want an agent to do X, so you give it the tools, a good system prompt, and if the model is smart enough, everything will be fine. And that's exactly what people use to prototype, and why wouldn't they? They're not wrong.
But here's the problem: it treats architecture as a consequence of the goal rather than something you design deliberately. Then the system grows. You know, more tools, more workflows, more edge cases, more users, and suddenly there's no real foundation under all of that.
The bottom-up approach takes more time, but it is much more comfortable. You start with the basic building blocks and make sure they actually work. Then you determine what each part needs to communicate, what data it owns, and what it is responsible for. Eventually, the system takes shape naturally from the interaction of its parts.
This isn't an argument of the "real engineers build everything from scratch" type. It's not even really a question of tools. It's a question of the mental model you build. Some engineers use sophisticated frameworks and build clean systems because they understand what each layer needs to do. Others write basic Python and create an undebuggable mess because they are still thinking in terms of "the agent decides everything." Tools flow from the model in your head, not the other way around.
One example of a particularly robust multi-agent system had almost no AI-specific infrastructure. At first glance, it looked like the wrong repository. A message queue, worker processes with distinct scopes, shared state storage with explicit read/write contracts, and a coordinator making routing decisions.
The language model queries were performed by the workers themselves, each receiving a set of context created upstream by another process. In total, it came to about a thousand lines of Python. Some demonstration agents have more code than that. Every part was traceable. When something behaved unexpectedly, the problem was usually found in under an hour because there was no magic to examine. Just code with a clear path through it.
This system was built bottom-up. The goal was defined, but the architecture was not derived from it. The components were designed first, evaluated individually, and then assembled to implement the desired functionality. That last aspect is the most important, not the first.
Where This Is Heading
The field is slowly shifting from "agent frameworks" toward proper infrastructure, with systems for evaluation, model routing, fallbacks, and state management. At least part of this already exists. The majority is still to come as people solve tough production problems in this space.
One observation comes up again and again: the people building the most reliable systems often don't even use the best models.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.