Uber and Microsoft Confront Soaring AI Token Costs

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Financial Pressure on Tech Companies
In the tech industry, there is growing concern about the costs associated with using artificial intelligence. Uber, for instance, has already exhausted its AI coding budget for 2026 by April. Microsoft, on its part, had to revoke the Claude Code licenses of its developers just a few months after they were activated. An employee from Priceline told TechCrunch that renewing the contract with Cursor cost four to five times more than expected.
Although the price per token has decreased, the increasing adoption of AI and increasingly autonomous agents has led to unprecedented token consumption. Companies that initially benefited from unlimited subscriptions now find themselves needing to understand where their money is going, cut their spending, and determine if they can achieve a return on investment from the remnants of their budgets.
A Changing Market
In the face of these challenges, a market is beginning to form to help companies manage these costs. Startups, established providers, and a new standardization body are striving to provide companies with the tools and language necessary to track their spending. Alexander Embiricos, head of business at OpenAI, stated at an event in New York that conversations with clients have evolved. "Six months ago, it was all about 'What can it do? Is it enough?' Now, it's 'Hey, we're spending so much. What visibility do you have? What auditability do you have? What token controls do you have? How efficient are your models?'"
The Tokenomics Foundation and Its Goals
It is in this context that the Linux Foundation unveiled plans for the Tokenomics Foundation, a new standardization body aimed at establishing the same cost discipline around AI tokens that FinOps has set for cloud spending. J.R. Storment, executive director of the FinOps Foundation, explained that by April and May, companies began expressing their concerns: "Oh my god, we are three times over our token budget for 2026 and it's only April." This situation has shifted the conversation from maximizing tokens and rapid innovation to the need for safeguards and control.
The Impact of New AI Models
The outcry heard in the tech world followed fervent demands from CEOs pushing their teams to use the best models and move quickly, regardless of cost. New models launched in November, such as Claude Opus 4.5 from Anthropic, GPT-5.1 from OpenAI, and Gemini 3 Pro from Google, brought significant improvements to agent tools, which multiplied consumption. One company found itself with a $500 million bill for Claude after forgetting to set usage limits for its employees.
Chris Reed, senior director of IT finance at Priceline, compared this situation to a "crack cocaine epidemic," noting that the company began imposing token limits on certain groups. Vitaly Gordon, CEO of the engineering operations platform Faros AI, mentioned that he recently spoke to a CTO who said, "One of my engineers spent $40,000 on tokens last month, and I really don’t know if I should stop him or tell everyone to do the same."
The Challenges of Measuring Productivity
A March survey conducted by Faros revealed that among 20,000 developers, output was increasing, but so were bugs and rewrites. Jellyfish, an engineering management platform, also found that engineers who used the most tokens were about twice as productive as those who used AI less, but they spent ten times more tokens to achieve that.
Nicholas Arcolano, head of research at Jellyfish, told TechCrunch via email that AI spending is exploding largely due to agent features, with consumption per developer increasing by about 18.6 times in nine months. Overall, these statistics make the case for productivity more ambiguous than spending suggests.
The Complexity of Cost Management
"Whether extreme spending is profitable depends on the ultimate business value of the shipped code (e.g., revenue), which most companies still cannot measure," Arcolano stated. Part of this measurement problem is the extent to which AI is used today. "Tracking cloud costs is a data problem of hundreds of millions of lines per month," Storment said. "Tracking token costs is a data problem of trillions of lines per month. You can't just put this into any spreadsheet or even basic tool. You need to fundamentally rethink your tools, specifications, and accounting systems to do it."
At Priceline, Reed is already noticing disparities. He noted issues between the usage reported by a provider and Priceline's internal data. "I started my career in telecom expense management, and I see all the same parallels, from telecom to cloud to AI," he said. "Every time you introduce something new, it's prone to billing errors and opportunities for audit and optimization."
An Expanding Market
A market is beginning to form around this issue. There are specialized companies, like Pay-i, which tracks, measures, and optimizes costs and performance of GenAI investments. Paid, on the other hand, allows developers to track costs, measure usage, and bill users based on actual value rather than subscription fees.
Then there are companies like Jellyfish, Waydev, and Faros AI, all providing monitoring of AI agents to prove the ROI of development tools. Storment indicates that most of the 180 providers within the FinOps Foundation are moving toward this space.
Companies with existing distribution are also adding new features to capitalize on this new market. Ramp has recently ventured into AI expense management; Datadog and New Relic have added services such as cloud cost management, token-level observability, and GPU monitoring. At the FinOps X conference next week, AWS is expected to introduce new financial management features focused on AI spending for businesses.
Towards Cost Optimization
Tiffany Luck, a partner at NEA, believes that token efficiency and observability will likely be added at the "harness or application layer." She cited Factory, a startup that creates AI agents for businesses, which launched a model router this week that automatically selects the right model for each task.
Gordon expects that leading labs and other model providers will adopt an OpenRouter-style optimization to direct queries to the cheapest models—a trend that is already appearing on corporate Claude bills. "The financial report on how much you spend on Anthropic, even if you call the model Opus, part of the spending will be on Sonnet or Haiku, because they are smart enough to do that," Gordon stated. "I think this will become increasingly common."
The Role of the Tokenomics Foundation
But all these tools are being built without a common language or shared definitions of how much a token costs, what it produces, and how to compare spending across providers. This is where the Tokenomics Foundation hopes to be helpful. The foundation is building a canonical definition and framework for "tokenomics;" open standards, specifications, and metrics for the use and billing of AI tokens; as well as new metrics for the AI economy, such as cost per intelligence or tokens per watt. It also plans to define metrics on the efficiency of token factories and consumption efficiency. The group anticipates a formal launch in July and is preparing to announce new members at the FinOps X conference next week.
Future Outlook
"The token economy is fundamentally more abstract and opaque than anything we've managed at this scale before," said Nishant Gupta, director of availability at Salesforce, in a statement. "This requires a different operational force than the industry has built for the cloud."
That said, Goldman Sachs predicts that global token usage will multiply by 24 by 2030. Companies already over their budget need solutions now, and the foundation's first deliverable is still months away. "Maybe we've created a steam engine, but we still haven't figured out the assembly line," Gordon said.
According to Arcolano, the smart move is broad and moderate adoption. "The best ROI comes from moving the broad middle from low to moderate usage, rather than pushing heavy users higher," he stated.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.