Brief IA

Google Frozen v2: the chip that challenges OpenAI and Anthropic

🔬 Research·Tom Levy·

Google Frozen v2: the chip that challenges OpenAI and Anthropic

Google Frozen v2: the chip that challenges OpenAI and Anthropic
Key Takeaways
1Google is developing the 'Frozen v2' chip with integrated Gemini architecture for its servers.
2It could be 6 to 10 times more efficient than current TPUs, according to internal sources.
3Scheduled for 2028, it could reduce AI inference costs and provide a competitive advantage.
💡Why it mattersThis chip could reposition Google as a market leader in AI against its major competitors.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Google Frozen v2: The Chip That Challenges OpenAI and Anthropic

Google's "Frozen v2" chip apparently integrates the architecture of the Gemini model directly into silicon for efficiency gains.

Google is developing a new server chip in-house, called "Frozen v2," which incorporates the architecture of the Gemini AI model directly into the silicon. According to sources cited by The Information, this chip could be 6 to 10 times more efficient at providing AI responses than Google's current TPU chips. Google plans to deploy it starting in 2028 and views Frozen v2 as a test for specialized chips, with a production volume smaller than its TPU range.

Unlike Google's TPUs, which work with many models, Frozen v2 has parts of the Gemini model structure integrated directly into the hardware. The name follows the same logic as "freezing" parameters in AI models, where you lock values so they stop changing. With Frozen v2, a portion of the model is permanently embedded in the chip itself, reducing computation steps and speeding up responses.

The original idea reportedly came from Jeff Dean, the chief scientist at Google DeepMind. His initial Frozen design aimed to directly integrate the model weights into the chip. Weights are the specific settings that determine how an AI model responds to queries. Google abandoned this approach because the chip would only work with a single version of Gemini and would become obsolete too quickly.

Frozen v2 takes a more flexible approach by integrating the model architecture rather than the weights, meaning the underlying plan rather than the adjusted parameters. New weights can still be loaded onto the chip. The amount of architecture that will actually be hard-coded has not yet been decided, according to The Information.

Optimizing Margins on AI Inference

Since the chip only functions as long as Google remains committed to the same model architecture, it is unlikely to become a product for external customers. Google is already leasing its TPUs to Meta, offering them to external cloud clients, and positioning them through its "TPU@Premises" program as an alternative to Nvidia, with an internal goal of capturing ten percent of Nvidia's annual revenue. In contrast, Frozen v2 is intended to alleviate internal pressure on Google's AI computing capacity.

If the chip lives up to its promises, it could nonetheless become a major competitive asset. In the AI sector, how companies optimize inference costs increasingly determines their margins. Google could use Frozen v2 to run powerful models at lower prices and capture market share from OpenAI and Anthropic.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.