AI: Inference and Context Impact the Bill

The cost of AI does not stop at training. At scale, the price per token, the length of context, the size of models, and the architecture of agents all affect the actual cost of usage. Four industrial layers share expenses and value differently, while a decrease in unit prices does not guarantee a lower total expenditure.
Decreasing Unit Costs, but Budgets That Can Rise
The cost of a given AI capability has decreased rapidly, driven by more efficient models, improved hardware, enhanced inference techniques, and competition that drives prices down. However, a lower unit cost does not necessarily mean a lower total expenditure. When a service becomes cheaper, its usage increases, a phenomenon observed for decades in computing. A model may cost ten times less to run, but if usage multiplies by twenty, the total computing expense still rises. In this context, the price of an API call is not the most relevant indicator: the right question is how much useful work is obtained per dollar invested. As AI moves from experimentation to production, understanding this economy is as important as understanding the models.
Who Pays and Who Profits: Four Layers, Four Realities
The AI value chain is organized into four layers with distinct cost structures. Chip manufacturers sell the required hardware regardless of the dominant model; it matters little, for example, whether OpenAI, Anthropic, or Google prevails as long as training continues. Cloud providers rent out data centers, networks, and GPU capacity, benefiting from demand without needing to build a winning model. Model companies absorb heavy costs for training, research, and inference infrastructure in a market that evolves every few months. Finally, the application layer relies on existing models to focus on specific problems without training a base model from scratch. In this landscape, the company that builds the product is not necessarily the one capturing the most value; the outcome depends on the layer considered and will likely evolve with prices and technology.
Inference Charges for Every Exchange: Tokens, Model Size, and Scale
Once a model is available, each use incurs a cost: this is inference. With each request, the model processes inputs, performs numerous calculations, and produces outputs, which explains the billing by token, roughly a small piece of text. More text in input and output means more computation, and larger models cost more per token, regardless of the question asked. Therefore, long conversations and large models quickly become expensive. At the scale of billions of requests, seemingly modest unit costs add up. Training is a one-time expense, while inference never really stops.
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
The Overlooked Cost: Context, Especially with Agents
Transitioning from simple chat applications to operational agents brings forth a context cost. An agent carries the history, system instructions, retrieved documents, tool results, database searches, and sometimes outputs from other agents, most of which are sent back to the model at each step. If it makes 10 calls while dragging a large context, one pays to reprocess the same information multiple times. A poorly designed agent carries unnecessary context at every step, while a well-designed agent retains only what is useful.
Reducing the Bill: Caching, Context Management, and Context Mesh
Several engineering decisions impact expenses. Caching prompts avoids reprocessing constant components, such as system instructions, tool definitions, or a reference document. Caching responses allows for returning a previously generated answer to similar questions. Context management involves deciding how many messages to retain, what level of detail of a tool call to pass on, or when to summarize a conversation sequence; at scale, these are also economic decisions. With MCP, agents benefit from a unified method to access tools, but in a structure involving several hundred agents, multiple and overlapping accesses to the same systems can occur: it is possible for two agents to access the same client file, request the same API, and each keep their own version. This type of duplication is a negligible detail at a small scale but becomes a cost and architectural issue when it becomes widespread. A context mesh, a shared layer between agents and tools, akin to an API gateway for services, interposes a unique mediation between agents and backends. The expected savings come from facilitated tool discovery and reduced duplicated work.
Deploying Internally: Limits, Architecture, and Useful Return per Euro
In a product, costs always manifest: usage limits exist because inference is not free, and caching prevents reprocessing the same thing twice. Context management matters because every unnecessary token is still charged, and the architecture of agents influences expenditure: an agent making ten unnecessary tool calls is slower and consumes more. For internal adoption, the challenge is whether AI can perform a task reliably, quickly, and economically compared to alternatives. Nothing is magical or free: behind the demonstrations lie infrastructure costs, real constraints, and competitive incentives.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.