AI Agents: Three Architectures to Limit Command Injection

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Recent incidents remind us that a single hijacking is enough to compromise an AI agent connected to the Internet and internal systems. Three architectural models—action selector, plan then execute, code then execute—offer concrete safeguards, provided one accepts compromises on flexibility and usability. Here’s how they are implemented and what they do not cover.
Effective Safeguards, but Reduced Adaptability
Each model studied provides targeted protection against command injection, at the cost of increased rigidity. The action selector can make an agent nearly immune to manipulation, but it limits the agent to a fixed set of operations and exposes it to implementation errors. A password reset agent built this way only executes predefined tasks and does not "design" anything beyond that.
The "plan then execute" model prevents an attacker from influencing the choice of tools or arguments, but it sacrifices real-time adaptability: the agent does not react to unexpected structures, missing information, or dynamic forms, and malicious content can still taint the output. Its "code then execute" variant locks the execution flow, without preventing manipulation of input variables. It also introduces risks inherent to code execution, which are sometimes unavoidable but must be handled with caution, even though, for certain tasks, writing and executing code proves more appropriate than simple planning.
Training the User and Locking Contact Points
Security starts with the system's endorsement: the user remains the weakest link, as Bruce Schneier reminds us. Even in good faith, internal employees can misinterpret a request, make a mistake, or unknowingly relay a hostile instruction via an external source.
Mapping possible actions, strengthening authentication and access checks, and providing comprehensive security training constitute the first line of defense. Connections to unreliable sources and downloads are identified as critical points: repositories are limited to internal documents selected from SharePoint, while recognizing that a simple copy-paste of web content can expose the agent.
Interposing Safe Services: The Example of Functions and Emails
Beyond instructions to the model, systemic defenses offer the best leverage for risk reduction. Instructions can filter out undesirable uses, but well-structured architectures effectively control the effects.
Specifically, the agent never executes raw SQL. It chooses query templates and fills in parameters, while APIs hosted via Azure Function ensure execution, limited to pre-programmed operations. A request for projects by sector illustrates this decoupling: the API responds after querying or triggers an error if the sector does not exist. The same principle governs email sending: a function checks that the recipient is on an authorized list before sending the message, without direct exposure to the messaging service.
Restricting Authorized Actions Reduces Injection Risks
Strictly limiting the repertoire of actions prevents the LLM from facing unreliable content and cuts short attempts at injection. In this framework, the model translates natural language queries into explicitly authorized operations within a known scope.
The benefit is particularly evident when an agent has access to the web and an internal database: if it could generate and execute SQL, a sneaky instruction inviting it to truncate tables could be followed to the letter. Even without high privileges on the database side, other vectors remain. By removing direct execution from the LLM and entrusting it to bounded components, the attack surface is reduced.
Planning or Coding Before Acting to Freeze Sensitive Decisions
The "plan then execute" model requires deciding in advance on the resources to consult and the email recipients before any exposure to unreliable sources. For an agent assembling a weekly report from internal data and web research, this discipline prevents a contact discovered online—potentially controlled by an attacker—from becoming a message recipient. While harmful content can still slip into the writing, the scheduling of tools and sensitive parameters, such as the sending address, are fixed and out of reach of a third party.
When the task structure requires a loop or logic whose size is unknown in advance—for example, unsubscribing from the last 100 newsletters never opened in 3 months, which may only be 3—"code then execute" transforms the request into code and locks the execution flow. The adversary can influence variable values but not the execution path, at the cost of risks inherent to code execution.
Why These Safeguards Are Necessary: Documented Failures
The perceived likelihood of an adverse scenario may seem low, but incident reports are accumulating, and a single one is enough to compromise a system. Cases have seen NotebookLM incorporate information belonging to another client into an image URL, the ChatGPT operator extract a private email from an authenticated Hacker News account, and Microsoft Copilot summarize a message by citing a link controlled by an attacker.
The environment is conducive to command injection: pages slip in implicit injunctions, models confuse context with instruction, and the range of outcomes goes from data leaks to destructive modifications in the database. Even without apparent malice, a competitor's page structured in a directive manner can skew rankings. This is the challenge to address upstream in the architecture.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.