AI Costs and Models: Challenges for Businesses in 2026
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
In the past two years, the rise of artificial intelligence (AI) in large companies has been meteoric. Every organization has launched at least one AI project, convinced by the promising results of pilot programs. Budgets have been approved, teams trained, and yet, when it comes time to scale, numerous obstacles arise.
The costs associated with AI are spiraling out of control. Models change without warning, and so-called "intelligent" agents often turn out to be mere automations. Contracts with major suppliers become golden traps. This article explores the realities on the ground, with figures, concrete cases, and the lessons that IT departments are beginning to draw.
A Bill Much Higher Than Expected
CFOs are surprised by the increase in budgets allocated to AI. In 2024, the average budget was $1.2 million per year, but it skyrocketed to $7 million by 2026, representing a 483% increase in two years, according to the "AI Inference Cost Crisis 2026" report published by Oplexa in March 2026.
This increase is paradoxical, as the unit cost of AI has dropped by 280 times in two years. This phenomenon is known as the "Jevons Paradox": when something becomes cheaper, its consumption increases to the point that total expenditure rises.
A technical explanation adds to this paradox. A standard chatbot makes a single call to the model to answer a question, while an autonomous agent, which plans, executes actions, checks its results, and self-corrects, makes between 10 and 20 calls to accomplish the same task.
Costs accumulate with each iteration. According to Gartner, more than 40% of agent-based projects will be canceled by 2027 due to these cost overruns. "Always-on" agents, which continuously monitor emails, systems, or market data, can generate monthly infrastructure costs exceeding $200,000 for a single application.
The initial cost of a pilot was €2,000, but scaling up costs two hundred times more. The platform Zylo, which tracks technology budgets for thousands of companies, reports that 78% of IT leaders surveyed were surprised by unanticipated charges.
The FinOps Foundation, in its "State of FinOps 2026" report, identifies AI as "the fastest-growing expense category" and warns that current API prices are partially subsidized by venture capital. A 30% to 50% increase in rates over the next 18 months is considered basic financial prudence.
Moreover, Chinese competition offers solutions up to 10 times cheaper, adding further pressure on Western companies to control their costs.
Unexpected Withdrawal of AI Models
On January 29, 2026, OpenAI announced the withdrawal of its models GPT-4o, GPT-4.1, and GPT-4.1 mini, with a deadline set for October 14, 2026, for API client companies. According to engineering firm TensorOps, GPT-4.1 was "the silent workhorse that powered half of the enterprise stack," valued for its predictability, economical cost, and stability, with an average latency of 1.35 seconds and a cost of $2 per million tokens in input.
This withdrawal is not a simple version update but a fundamental architectural change. The replacement models "reason" in multiple steps before responding, which increases the average latency of the GPT-5 API to 4.26 seconds, with peaks measured at 21.7 seconds, and the cost of output tokens has risen by 25%.
For applications where user experience depends on responsiveness, such as an assistant in a call center or a commercial co-pilot, these figures change everything. Microsoft began automatically upgrading Azure OpenAI deployments to GPT-5.1 starting March 9, 2026, with an official end-of-life for older versions set for March 31, 2026.
Organizations that had not planned for this migration discovered the change afterward, in their performance metrics, in their costs, or in unexpected error alerts. This is not the first time OpenAI has deprecated models, such as GPT-3.5-turbo-0301, GPT-4-0314, and GPT-4-0613 since 2023.
For companies whose architecture is directly tied to a specific model version, each deprecation is a forced migration, involving weeks of engineering, regression testing, and business validations restarted.
The lesson is simple to articulate but difficult to implement retrospectively: a robust AI architecture cannot depend on a single model. It must treat models as interchangeable components, like light bulbs in an electrical installation, and not as the wiring itself. This is precisely where the use of open-source models deployed on dedicated GPU infrastructures makes perfect sense: stability, autonomy, and protection against the volatility of proprietary APIs.
Data from 2026 shows that this approach can reduce inference costs by 60% to 80% compared to high-volume proprietary APIs.
The Trap of Major Platforms
In the face of model disappearances, the temptation is to entrust the entire AI infrastructure to one of the major cloud providers, which offer managed environments, service guarantees, and model catalogs.
However, the risk of "vendor lock-in" is very real. A survey published in February 2026 by Parallels, conducted among 540 IT professionals in the United States, the United Kingdom, and Germany, reveals that 49% are considering bringing part of their infrastructure back to hybrid or on-premises environments. And 84% are now integrating digital sovereignty into their AI strategy.
The paradox is that only 6% of organizations are able to change providers without significant consequences. The remaining 94% know they are trapped but do not know how to escape.
This risk is not theoretical. In 2025, Builder.ai, a startup valued at $1.3 billion and backed by Microsoft, abruptly ceased operations. Clients who had built critical processes on its platform found themselves without access to their systems, data, and workflows. NexGen Manufacturing, one of the documented cases, spent $315,000 to migrate its 40 AI workflows to a new platform.
The firm Andreessen Horowitz articulates the structural lesson as follows: "The rise of agentic workflows makes changing models increasingly difficult. The more invested the guardrails and prompts are, the more organizations hesitate to change."
Contract law figures confirm that the distrust is justified. A study by Stanford Law School published in 2025 analyzed the terms and conditions of major AI providers: 92% claim broad rights over the data that passes through their systems. Only 17% commit to full regulatory compliance. And only 33% cover risks related to the intellectual property of generated content.
The European AI Act provides for penalties of up to €35 million or 7% of global revenue.
Claude, Microsoft, Fable 5, and Kimi: A Revelatory Week
To illustrate the instability of the AI model market, recent events are telling. Microsoft canceled Claude Code due to the bill.
On May 14, 2026, The Verge revealed that Microsoft was canceling almost all of its Claude Code licenses. The Experiences & Devices division, which produces Windows, Microsoft 365, Teams, and Surface, had been one of the first to deploy Claude Code in December 2025, allowing thousands of its own developers to use the Anthropic tool daily. The verdict, six months later, is that the consumption-based token billing has made the expense untenable. The licenses are terminated as of June 30, 2026, and developers are switching to GitHub Copilot CLI.
The wording of the internal memo, signed by Executive Vice President Rajesh Jha, is diplomatic: "Our goal was to learn quickly, to compare tools in real engineering workflows."
The reality is simpler: even Microsoft, one of the largest investors in generative AI with over $13 billion injected into OpenAI, cannot absorb the costs of intensive use of Claude Code at the scale of a large organization.
The lesson is not that Claude is bad. It is that token billing, at scale, is a business model that even the largest companies in the world do not master.
The U.S. government cut Fable 5 within 48 hours.
On June 12, 2026, the U.S. government ordered Anthropic to cut access to its two most powerful models. Claude Fable 5 and Mythos 5 were taken offline at 5:21 PM Eastern Time, based on an export control directive citing national security reasons.
In its statement, Anthropic writes: "Our understanding is that the government believes it has identified a method of circumvention, or 'jailbreak,' of Fable 5. The letter received did not provide any specific details about the national security threat." Anthropic suspended access for all its clients.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.