Generative AI: Soaring Costs Worry Tech Giants

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Generative AI: A Cost Explosion That Worries Tech Giants
Generative artificial intelligence (AI) was expected to become cheaper as it developed. Instead, the opposite is happening, and the bill is starting to weigh heavily, even for tech giants. AI is not a success in every respect, far from it.
For years, AI companies have sold the same promise to investors: the larger the model, the more powerful it becomes. This is known as “scaling laws.” However, this race for size is starting to become very expensive. For example, Uber exhausted its entire budget for AI in 2026 in just four months. Microsoft has pulled some Claude Code licenses from its employees, deeming them too costly, and has canceled several data center projects. Atlassian, Adobe, and Amazon have also tightened their belts. These setbacks are not isolated incidents but reveal a much deeper problem, as indicated by a comprehensive survey from The Atlantic.
A Scaling Problem
When a streaming service or mobile app gains millions of additional users, the cost of each new user decreases. This is known as economies of scale, and it’s exactly what investors are looking for: the larger the company grows, the more profitable it becomes.
However, generative AIs like ChatGPT or Claude operate in the opposite direction. The more text they are asked to process, the more memory and computing time they require. This increase is not proportional, as it accelerates with rising volumes. Thus, doubling the amount of text to analyze can lead to a bill that increases by much more than double.
This technical flaw directly impacts the budgets of client companies, as most contracts charge for AI usage based on the volume processed, the infamous token. The result: the bill rises with the intensity of use, without the actual value produced for the company keeping pace.
The Impacts Are Already Concrete
This inefficiency has colossal repercussions. To run their models, AI giants are now purchasing up to 70% of the global supply of high-end computing memory. The direct consequence is that prices for hard drives, computers, and smartphones are skyrocketing, and the trend shows no signs of stopping. Moreover, the energy impact of this dynamic is disastrous. To meet such demand, the capacity of American data centers could be multiplied by eight in the coming years, to the point where some companies are considering recycling airplane engines to produce enough electricity. This represents an even more challenging hurdle as advancements in hardware slow down: components have reached physical limits that hinder their miniaturization.
In this context, the profitability of the sector remains a huge question mark, fueling fears of a speculative bubble around AI, a sector in which entire segments of the economy have massively invested.
To Discover
- What are the 5 best AI chatbots? Comparison 2026
Frequently Asked Questions
What exactly do the “scaling laws” in generative AI cover?
Scaling laws describe an empirical relationship between three variables: the size of the model (number of parameters), the amount of training data, and the computing power used. In practice, they indicate that as these resources increase, performance improves quite predictably on certain indicators (perplexity, benchmark scores). The key point is that the gain is often marginal: each “performance tier” costs significantly more than the previous one. They mainly concern training but also push for deploying larger models, which are therefore more expensive to run. Finally, they do not guarantee that gains will automatically translate into business value or cost reductions.
Why can processing more text cost disproportionately more with a large language model?
When a model generates or analyzes text, it must process a sequence of tokens and maintain a “context” in memory to produce a coherent response. In the Transformer architecture used by most large models, the attention mechanism has a cost that tends to grow quadratically with the length of the context (simplifying: the more tokens there are, the more relationships need to be calculated). This increases both the computing time (GPU) and the amount of memory required, which directly impacts the bill. Even when optimizations exist (more efficient attention, context compression), they do not eliminate the problem for very high-volume uses. The result: lengthening prompts, multiplying documents to ingest, or increasing context windows rapidly drives up costs.
What is token billing, and why can this model explode a company's budget?
A token is a unit of text used to measure the input (prompt) and output (response) of a model; it is not exactly a word, but rather a piece of a word or punctuation. Most AI APIs charge separately for incoming and outgoing tokens, sometimes with a higher rate for more powerful models or larger context windows. This billing method directly ties costs to usage intensity: the more automation, the more documents ingested, the more text generated, the higher the expenses mechanically rise. It can also penalize cases where the value produced does not grow at the same rate as the volume of tokens (e.g., long responses, reformulations, iterations). Without safeguards (quotas, routing to smaller models, response caching), rapid adoption can exceed projected budgets.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.