⚡
Brief IA
›

Corporate Agents: Human Oversight and Context Management

🔬 Research·Tom Levy·

Corporate Agents: Human Oversight and Context Management

Corporate Agents: Human Oversight and Context Management
⚡
Key Takeaways
1Enterprise applications require explicit human oversight for sensitive actions of AI agents
2Context engineering, prompt caching, and pruning observations are essential for managing costs and latency
3The ReAct model offers flexibility and error recovery but remains costly and prone to drift on complex tasks
💡Why it matters — These safeguards and optimizations are crucial for making AI agents reliable and economically viable in production.

Deploying AI agents in production is not merely a matter of an autonomous loop. Human supervision, delegated decisions, meticulous context management, and caching are crucial for security and cost management. The ReAct model serves as a foundation, but its promises of flexibility come at the cost of latency and tokens.

Supervise Before Executing: The Essential HITL Step

Enterprise applications do not delegate sensitive actions to agents without oversight. Given that language models are non-deterministic, any operation beyond reading can become problematic, such as deleting a SQL table, authorizing a payment, or sending an email to a customer. The Human-in-the-Loop mechanism explicitly interrupts execution and awaits review. In a state graph implementation, for example with LangGraph, invoking a specific tool can halt the graph. The state is serialized in a database, and execution is suspended. An examiner then reviews the output (for example, a drafted email), can edit it, and then approves it for the orchestrator to resume with the validated state.

Cost and Latency: Caching and Pruning as Levers

Context Engineering is central to production. An agent's context is constantly evolving, accumulating observations, errors, and intermediate thoughts. Systematically loading this data inflates the context window, increases latency and costs, and can cause the model to lose fine details, resulting in inaccurate responses for downstream agents. Providers offer prompt caching: system instructions and tool definitions, which can exceed 5,000 tokens, are placed at the top of the context to reuse cached key-value attention states. During repeated loops, only new observations are charged, which can significantly reduce inference costs. At the query scale, pruning observations prevents a large document (like an HTML page) from saturating the window. Compressing past steps, for example by summarizing the last 5 search results with a smaller and less costly model, adds to these gains.

Cheaper Decision-Making: Offloading Determinism from the LLM

An agentic system must frequently decide on the next steps in a workflow: next step, delegation to an agent, tool to mobilize, sufficiency of context. Calibrated and low-cost decision models, such as TypeSafe’s JEV, handle some of these deterministic choices. This approach aims to leave reasoning and synthesis tasks to language models. Meanwhile, architectures are moving away from single-agent loops towards constrained multi-agent flows, following developments like specialized deterministic topologies illustrated by GraphRAG.

⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

Four Pillars to Assemble: Model, Tools, Control, and Memory

An agent relies on a base LLM, tools, a planning and control mechanism, and memory. The language model serves as the reasoning engine and benefits from natively supporting tool or function calls, with examples cited such as GPT-5, Claude 5.5 Sonnet, or Llama-4. Achievable tasks include using API wrappers (search_web, query_database, send_email) as well as Python interpreters operating in an isolated environment. Control is exercised either strictly, through a state machine or a DAG, or flexibly via an open loop. Memory combines a short-term prompt window and vector stores for past episodes. On the context side, state projection removes from the prompt what is no longer useful, such as search logs when transitioning to drafting. Structured notes in JSON or XML make reasoning filterable from turn to turn, while semantic retrieval only brings back relevant historical actions.

ReAct: Functioning, Gains, and Operational Limits

The ReAct model explicitly alternates between reflection and action, chaining tool calls and observations until a final response is reached. The agent starts from a system prompt that exposes its personality, tools, and rules, then loops at each iteration with a thought and an action, with orchestration executing the tool and returning the observation to the model. The context grows linearly over turns. Pseudo-code illustrates a loop limited to 10 iterations, summarizing observations beyond 2000 characters, returning tool errors as observations, and stopping with a message if the task fails. Strengths include flexibility and error recovery, for example after a 404 API response. The trade-offs are latency and cost, as a request may require 10 sequential calls, with notable token consumption even with caching, and risks of drift or hallucination on complex tasks. ReAct is suitable for open-ended research-oriented tasks, such as data analysis, in-depth web research, or sandbox debugging, and is generally discouraged for deterministic low-latency conversations requiring precise responses.

When One Agent is No Longer Enough: Sequencing and Specializing

Beyond a certain level of complexity, entrusting an entire task to a single ReAct loop is unreliable. Decomposing into subtasks and processing them through a sequential multi-agent workflow allows each step to be assigned to a specialized agent.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.