Brief IA

Nvidia and the Token Dilemma: Optimizing Without Sacrificing Humanity

🤖 Models & LLM·Tom Levy·

Nvidia and the Token Dilemma: Optimizing Without Sacrificing Humanity

Nvidia and the Token Dilemma: Optimizing Without Sacrificing Humanity
Key Takeaways
1Jensen Huang of Nvidia emphasizes the importance of a high token budget to justify the cost of engineers.
2Large companies are investing heavily in AI, but layoffs do not guarantee a return on investment.
3Strategies like prompt caching and the use of smaller models can significantly reduce token costs.
💡Why it mattersEffective management of token budgets could transform how companies balance technological innovation and job retention.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Nvidia and the Strategic Management of Tokens

During a notable appearance on the All-In podcast, Jensen Huang, the CEO of Nvidia, addressed a crucial question for the future of the tech industry: how to justify the cost of engineers relative to their token consumption. According to him, if an engineer with an annual salary of $500,000 consumes less than half of that amount in tokens, it should be a cause for concern. Nvidia has set an ambitious goal: an annual budget of $2 billion in tokens for its engineering teams.

This statement highlights a compromise that many companies have already discreetly adopted: redirecting funds traditionally allocated to salaries towards the purchase of tokens. The four largest hyperscalers plan to invest around $700 billion in capital expenditures by 2026, nearly double that of the previous year. Meanwhile, data from Challenger, Gray & Christmas reveals that AI has been the primary reason for layoffs in the United States for four consecutive months.

Layoffs as a Financing Strategy

An internal memo from Meta, revealed by Reuters, detailed the elimination of 8,000 positions in May, justified by significant investments, even as the company's revenue increased by 33% that quarter. These layoffs are not merely survival measures but a way to finance innovation.

However, these funding efforts have not always produced the expected results. A survey by Gartner, conducted among 350 executives from companies generating over $1 billion in revenue and using AI, showed that about 80% of them had reduced their workforce without any notable improvement in returns. Helen Poitevin, an analyst at Gartner, was clear: “Staff reductions can free up funds, but they do not improve outcomes.”

Uber and Token Management

Uber has experienced the costly management of tokens by providing 5,000 engineers with AI coding tools in December, exhausting its AI budget for 2026 by April. Although 70% of the code produced is generated by AI, Chief Operating Officer Andrew Macdonald admitted that the link to customer experience remains to be established: “That link is not there yet.”

These examples show that companies have often viewed token spending as fixed and labor as variable, while the reality may be the opposite. Layoffs lead to a loss of institutional knowledge, while the token budget can be adjusted with targeted efforts.

Optimizing the Token Budget

One of the most effective and economical solutions is to avoid paying multiple times for processing the same text. Prompt caching, now standard among major API providers, can reduce costs for repeated inputs by up to 90%. This is possible because static content, such as system instructions and reference documents, is processed once and then reused at a lower cost.

ProjectDiscovery, a security company, documented an increase in its cache success rate from 7% to 84% by restructuring its prompts, thereby reducing its total LLM expenses by 59 to 70% while serving 9.8 billion tokens from the cache. This approach has allowed for budget recovery greater than most AI-related layoffs.

Choosing the Right Model for Each Task

Another lever is to assign tasks to the appropriately sized model. Flagship models cost five times more than their smaller counterparts per token, yet many workloads continue to default to the most expensive models for classification and summarization tasks. Batch processing offers an additional 50% discount for tasks that do not require real-time responses.

Augmented generation by retrieval offers another approach by sending only the relevant part of a knowledge base to the model, rather than the whole. Prompt compression also reduces redundant examples that weigh down each call. Open-weight models can further reduce costs, managing routine workloads at a fraction of the prices of leading APIs for teams willing to manage the infrastructure.

Budget Discipline and the Future of AI

These measures are the AI equivalent of turning off the lights in unoccupied rooms. Uber has imposed a cap of $1,500 per month per engineer after exceeding its budget in April, proving that budget discipline is beginning to take hold. Companies anticipating these adjustments choose to act before the budget forces them to.

Investing in People: A Necessity

Optimizing token spending only makes sense if the savings are reinvested productively, and the strongest evidence indicates that this investment should be directed towards people. According to Helen Poitevin, organizations that have improved their return on investment are those that use AI to amplify their workforce rather than replace it.

Klarna conducted a revealing experiment by replacing about 700 customer service positions with an AI assistant powered by OpenAI, only to find a decline in customer satisfaction. CEO Sebastian Siemiatkowski admitted to Bloomberg: “The outcome was of lower quality, and that’s not sustainable.”

Today, Klarna operates on a hybrid model where AI handles routine volume, while re-hired humans manage tasks requiring judgment. Gartner predicts that by 2027, half of the companies that reduced their customer service staff for AI will rehire them.

The Challenge of Employing Young Developers

Investment in the workforce has become urgent rather than optional. The Institute for Human-Centered AI at Stanford University found that employment for software developers aged 22 to 25 has dropped by nearly 20% compared to 2024 levels, while older cohorts have increased. This means companies are eliminating the training ground for senior engineers they will need to lead these systems in five years.

A company that successfully reduces its token bill by 60% has the budgetary leeway to continue hiring at the base level. If it does so, it is a leadership decision, not merely a financial one.

Jensen Huang's reflections from Nvidia will continue to resonate during earnings meetings, and capital expenditures will keep rising. The companies that succeed will not be those that spent the most on tokens or laid off the most people to finance them, but those that understood from the start that the token budget was flexible, optimized it through engineering rather than workforce reductions, and invested the difference in the people who make those tokens valuable.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.