LLM: Cache Costs 10 Times Less, Avoid Double Toll

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Price discrepancies of 10x separate models within the same range, and cached tokens are worth one-tenth of normal tokens. Between variable cache timers and replayed history, part of the AI budgets melts away by spring. Usage rules and tiered model selection help contain costs without sacrificing results.
Delegate to Sub-Agents Without Losing Cache Savings
If the main agent operates on the Anthropic API, its cache window is 5 minutes. A sub-agent that takes 7 minutes to respond expires this cache and forces the manager to recompute all its context, with a cost that can exceed the expected savings. It is therefore recommended to delegate only tasks that reliably finish within the cache window. Sub-agents have a reputation for burning budgets, and asking a thousand agents to work on a slogan can cost a thousand times more than a single agent. However, when used correctly, they can reduce costs: treating the main agent as a senior manager, assigning clear tasks to sub-agents with just enough context—summarizing a file, making a yes/no decision, searching a folder or database and reporting the relevant information—allows these steps to be executed on cheaper models while keeping the manager's context focused and light.
Pricing: Higher Outputs and 10x Discrepancy Between Models
Each API call charges for incoming and outgoing tokens, with the output being the most expensive. An example cited by Anthropic: Claude Fable costs $5 to $10 per million tokens in input and $50 per million in output. Within this range, Fable costs double that of Opus, which is itself five times more expensive than Haiku, resulting in a 10x gap between the cheapest and the most expensive model. These rates fluctuate and require verification with the provider. Conversely, cached tokens are charged at one-tenth of the normal price, a major lever for reducing the bill.
Why History Increases Costs
Retaining the calculations of an exchange costs the provider, which limits the caching duration with a timer specific to each tool: 1 hour for subscription-based Claude apps, 5 minutes for the Anthropic API, about 5 to 10 minutes for ChatGPT, and 30 minutes for the OpenAI API. Without caching, the model replays the entire conversation at each turn, and you pay for the history both in input and output. Caching stores these calculations and, when warm, avoids recomputations, only processing the new part. Each use resets the timer, and the conversation can remain warm indefinitely as long as it progresses, for example, all afternoon within a one-hour window. Conversely, opening multiple tabs and returning hours later leaves the history visible but loses the cache, forcing an expensive complete recalculation. Asking one question per day in the same thread can thus cost several times more than grouping seven questions spaced five minutes apart.
Stay Within the Window, Write Concisely, Limit Context and Tools
Identifying your cache window and making your back-and-forth within it avoids recomputations. It is essential not to let an unfinished task exceed this window. Because the entire conversation is replayed when the cache is missed, every unnecessary phrase adds recurring weight; a concise session costs less than a verbose one. Reducing your messages, asking for brief responses, keeping instructions and prompts short, and requesting outputs in Simplified Technical English (ASD-STE100 standard) help contain tokens. Overloading the context degrades quality and increases the bill; it’s better to provide only what is essential. Finally, activating too many tools lengthens the system prompt and multiplies the risks of unnecessary calls. When the task changes, it’s preferable to open a new chat: two short, targeted exchanges are more economical and often more effective than a long, disjointed thread.
Choose the Right Tier and Pay for Reasoning Only When It’s Useful
It is not recommended to send all tasks to the most expensive model. Providers offer tiers: at Anthropic, Haiku, Sonnet, Opus, Fable; at OpenAI, the GPT-5.6 family with Luna, Terra, Sol. For light workloads (sorting, labeling, extraction, summaries), models like Haiku or Luna are sufficient; for mid-range tasks (code, documents, light analyses), Sonnet or Terra; for in-depth analysis or agent coordination, Opus, Fable, or Sol. Lightweight models do not always reason by default; those that do write a private chain of thought, which has supported recent advances but adds cost. For labeling 500 emails, this reasoning is unnecessary. Taking two seconds to gauge the difficulty before hitting Enter helps choose the right model, extending the subscription or API budget. Not all tasks need a high-end model.
Why Budgets Melt Away: Agents, Steps, and Fluctuating Prices
Many companies have already excessively tapped into their AI budget, some having depleted it as early as April, sometimes within a single quarter. This situation is explained by the frequency of use, but especially by the transformation of usage: there is a shift from chatbots to agents that successively use tools, read files, and perform checks, each phase involving token consumption. A request made to an agent can represent a cost several times higher than that of a simple interaction. Prices, meanwhile, are evolving: subscriptions, on-demand APIs, occasional promotional offers, and the arrival of new premium models are multiplying, complicating tracking. Understanding precisely what is being charged—including that a word is not a token, that you can count about 1.33 tokens per word, and that segmentations vary by model—becomes a prerequisite for regaining control over spending.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.