Brief IA

Amazon Reduces Dependence on Anthropic to Cut Alexa Costs

🔬 Research·Tom Levy·

Amazon Reduces Dependence on Anthropic to Cut Alexa Costs

Amazon Reduces Dependence on Anthropic to Cut Alexa Costs
Key Takeaways
1Amazon is revising the architecture of Alexa to reduce its dependence on expensive models from Anthropic.
2AWS cloud costs for Alexa+ could triple, reaching $1.7 billion by 2026.
3The company is focusing on its own AI models and better utilization of GPUs to optimize costs.
💡Why it mattersThis strategy could transform how companies manage the costs of AI services, influencing the overall AI market.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Amazon Seeks to Cut Alexa Costs by Reducing Dependence on Anthropic

In an effort to decrease the expenses associated with the artificial intelligence of its voice assistant Alexa+, Amazon is working to lessen its reliance on Anthropic's Claude models. The company is also exploring ways to optimize GPU usage to make operating Alexa+ more economical. Internal projections at Amazon indicate that AWS cloud costs for Alexa+ could nearly triple this year.

Amazon has begun restructuring Alexa to reduce its dependence on Anthropic's models as part of a broader strategy to cut operational costs for its AI-based voice assistant. According to internal documents reviewed by Business Insider, these changes have been implemented since late last year through early this year. The goal is to direct more queries to Amazon's internal AI models, thereby reducing unnecessary calls to Anthropic's Claude models and maximizing the efficiency of each GPU.

These initiatives are expected to quadruple the number of customer transactions that each unit of computing capacity can handle. This illustrates the new battleground of AI, where competition is shifting from creating smarter AI to reducing operational costs. Companies like Google, with its lower-cost AI Gemini Flash, and OpenAI, with automatic routing techniques to less expensive models, are following this trend.

Amazon's Financial Challenges with Alexa+

Amazon's financial forecasts highlight the reasons why the company is dedicating so much effort to this challenge. Early-year internal projections revealed that AWS cloud costs for the enhanced, AI-powered Alexa+ could reach around $1.7 billion by 2026, nearly triple that of the previous year. Alexa+ was also expected to operate about 60% above Amazon's target for AWS cloud costs per monthly active user. Despite identifying around $450 million in potential savings, internal reviews concluded that the company would not meet its financial goals. Amazon declined to comment on these projections.

A New Alexa More Expensive to Operate

Unlike previous versions, Alexa+ generates many responses using large language models running on GPU-intensive cloud services. This has transformed relatively low-cost voice requests into far more expensive AI workloads to process. These costs have escalated as Amazon faced a challenging launch. Business Insider previously reported that the company had delayed Alexa+ multiple times due to AI hallucination issues and questions about the service's readiness for customers. Alexa+ expanded its availability in the U.S. earlier this year.

The expansion of the service has increased financial pressure, highlighting the difference between generative AI and more traditional software services. As Alexa+ was deployed to more users, Amazon projected a net increase in AWS cloud expenses as demand for AI computing capacity grew. The company even considered delaying some of its most costly AI initiatives. Business Insider previously reported that Project Moonraker, Amazon's effort to equip Alexa with more advanced AI agent capabilities, was set to become the largest AI expense for the service this year, and the company considered delaying parts of the project to achieve savings.

Reducing Unnecessary Calls to Claude Models

One of Amazon's priorities was to reduce the use of Anthropic's Claude models within Alexa+. Internal roadmaps planned to shift specialized Alexa "Experts" from Claude Sonnet to Amazon's own AI models while reducing the use of other Claude models in the digital assistant service. Amazon also sought to avoid inference whenever possible. Inference is how AI models are executed, and one way to limit the cost of this is to use caching, which stores responses to common requests so that the AI does not have to redo the same work.

An Amazon roadmap projected that Alexa+ would stop calling Claude models when appropriate responses were already available in cache and expand "deterministic" processing, which allows Alexa to respond to more predictable requests without resorting to a large language model. This strategy is notable given Amazon's close ties with Anthropic. Amazon has invested billions in the AI startup, collaborates closely with it, and could reap significant benefits if Anthropic's IPO materializes.

However, internal documents reviewed by Business Insider show that Amazon is seeking ways to reduce the frequency with which Alexa relies on Anthropic's models. Amazon's approach reflects a growing trend in the AI industry. Investment firm William Blair wrote in a recent report that software companies are beginning to reserve cutting-edge models for difficult and high-risk reasoning while directing less complex requests to cheaper models. This reduces inference costs without changing the customer experience. "Multi-model routing is becoming a standard architecture in software," analysts at William Blair wrote in the report.

Optimizing GPU Usage

Reducing model costs was only part of the strategy. Amazon also focused on increasing the amount of work each GPU could perform. Rather than simply adding more Nvidia GPUs, Amazon aimed to handle more customer requests from the same hardware. A roadmap projected that software updates would increase available computing capacity by about 50% while reducing response times by approximately 40%. Internal planning charts tracked projected customer growth, GPU usage, available capacity, and inference efficiency as Amazon prepared to scale Alexa+.

Amazon's cost-saving efforts extended beyond software. Planning documents show that the company was evaluating both Nvidia GPUs and its own Trainium chips to further reduce the operating cost of Alexa+. More broadly, the documents indicate that Amazon views cutting-edge AI models and GPU capacity as costly resources to deploy selectively rather than by default. This philosophy echoes a point that CEO Andy Jassy has publicly expressed. In his letter to shareholders last year, Jassy argued that there is an "urgency" to make AI inference significantly less expensive.

"Reducing the cost per unit in AI will allow AI to be used as widely as customers desire, and will also lead to higher overall spending on AI," Jassy wrote.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.