Brief IA

Modelmaxxing: Optimizing AI to Cut Costs

🤖 Models & LLM·Tom Levy·

Modelmaxxing: Optimizing AI to Cut Costs

Modelmaxxing: Optimizing AI to Cut Costs
Key Takeaways
1Morgan Linton from Bold Metrics optimizes the use of AI models to reduce costs without sacrificing quality.
2Tokenmaxxing, once popular, is being replaced by a more targeted and cost-effective approach in businesses.
3Model routing tools are emerging to help companies choose the right AI model for the task.
💡Why it mattersCompanies are looking to maximize AI efficiency while controlling expenses, thereby transforming their technological approach.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Morgan Linton's Strategy at Bold Metrics

Morgan Linton, the Chief Technology Officer of the AI startup Bold Metrics, has taken steps to optimize the use of artificial intelligence models within his team. By orchestrating the use of different models according to specific needs, he decided to assign the use of Claude Fable, a low-cost model, to one team, while another focuses on leveraging GPT-5.5, a more advanced and expensive model. A third team, on the other hand, uses Cursor with Composer 2.5, achieving results he describes as "totally perfect." This approach allows Linton to avoid imposing strict limits on tokens while ensuring efficient use of available resources.

Being precise about model usage enables Linton to maximize efficiency without having to set strict caps on tokens. "My team uses the best resources, but in a much more efficient way," he stated.

The End of Tokenmaxxing

In 2026, the term "tokenmaxxing" dominated discussions in the AI field. This concept encouraged companies to maximize the use of AI models by their employees. However, after analyzing the associated costs, giants like Uber and Microsoft have adopted a more thoughtful approach. Rather than focusing on maximum usage, they now prioritize model switching. This strategy involves reserving the most expensive and high-performing models for complex tasks while using older, economical models for simpler tasks.

Founders, software engineers, UX designers, and even non-technical enthusiasts are discovering a way to save: model switching. They assign their most challenging and intellectually demanding tasks to more expensive, cutting-edge models while delegating simpler, repetitive tasks to older, cheaper models.

Resource Optimization

Kaylin Voss from OpenAI emphasized the importance of using recent models to reduce retries and unnecessary efforts. However, some tasks do not justify the use of expensive models. Brian Armstrong, CEO of Coinbase, predicted that 80% of workloads could be handled by much cheaper models in the near future, while the remaining 20% would require the latest models for optimal performance.

Chris Maconi, co-founder of Hechura, has always been skeptical of tokenmaxxing. He favors an approach where humans remain at the center, avoiding reliance solely on bots for coding. Maconi recalls the hype cycle around OpenClaw, an AI agent encapsulated in a Mac Mini, which was particularly token-hungry due to its 24/7 usage and autonomy. When he set up his OpenClaw, Maconi started with cheap Gemini models before transitioning to Anthropic's Haiku. "I'm not afraid to try some of these lower-performing models to see if they can provide the intelligence we need," he said.

Maximizing Token Efficiency

Tanvi Pisal, a UX designer, has learned to use models more strategically. By using Figma to design her projects before submitting them to Claude, she has managed to save tokens. She has a business subscription to ChatGPT and pays for the Claude Pro package at $20/month. Initially, she used Claude to brainstorm UX from scratch, a process that caused her to "burn through months of tokens" without completing the task.

"Now, I design everything in Figma first, then I put those screenshots into Claude. I tell Claude to keep the user interface as is and to build all the functionality and flow," Pisal added. "Doing this design process first really helps me save tokens."

She also chooses to brainstorm ideas with ChatGPT — which is free for her thanks to her business plan — then takes the refined ideas to submit to Claude for creating more polished documents.

Alejandra Thomas, a software engineer and tech content creator in New York, tests every new model released to see what each is good for. "I try not to use the most expensive or advanced model just because it's available. For simple tasks, I still use lighter models or none at all," Thomas stated.

Ed Stevens, CEO of the AI sales company Scoot, enjoys "picking a model and sticking with it." His engineers choose a model, try it for a few months, then determine if it meets their needs. If there's a shiny new model — or if they think they can achieve the same result for less — they switch models, Stevens explained.

The idea of maximizing the use of each token illustrates a scarcity mindset, according to Dan Ariely, a behavioral economics researcher and professor at Duke University. Ariely noted that token budgets remind him of old phones, which came with a limited number of calling minutes. People would try to maximize their minutes at the end of the month, even if it meant calling people they didn't really want to contact.

"Tokens create a scarcity model where people can't use as much as they want. This creates a usage goal and a psychology of waste if people don't meet their goal," he added. He also noted that, not wanting to exceed the limit and incur extra charges per use, users switch to models from other companies to save money once they hit their token ceiling.

The Rise of Model Routing Tools

For those who find modelmaxxing exhausting, startups specializing in model routing offer a solution. These companies, like OpenRouter, develop software that assigns tasks to specific models based on their complexity. David Gilmore leads one of these companies, Rayline. His tool intercepts requests and determines if they can be directed to cheaper models, often open-source. Many of his clients fall into the "FOMO moment" trap, he said. Then they receive their API bill and realize they need to cut back on spending.

The number of companies using a routing platform is slowly increasing. Ara Kharazian, Chief Economist at Ramp, told Business Insider that about 1% of companies used a model router last year; this year, it's 5%.

The San Francisco-based investment firm BlockSpaceForce uses OpenRouter, Fireworks, and Together AI. Spencer Yang, its managing partner, has also advocated for first asking a cheaper model if a more expensive one would be necessary for your task. "Models themselves are becoming really good at assessing their own complexity," Yang stated.

Some companies continue to default to the latest and most expensive models. Maconi, co-founder of Hecura, attributes this to laziness. "People don't want to make the effort to understand which models are good for which tasks," he said. "They just want to follow the trend."

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.