Brief IA

Grok 4.20: a Reliable AI Model but Outpaced by Gemini and GPT-5.4

🤖 Models & LLM·Tom Levy·

Grok 4.20: a Reliable AI Model but Outpaced by Gemini and GPT-5.4

Grok 4.20: a Reliable AI Model but Outpaced by Gemini and GPT-5.4
Key Takeaways
1Grok 4.20 from xAI scores 48 on the Intelligence Index, far behind Gemini 3.1 Pro Preview and GPT-5.4, which reach 57.
2Despite its lower performance, Grok 4.20 stands out with a record non-hallucination rate of 78% in the AA Omniscience test.
3The model offers three API variants and a context window of 2 million tokens, with a competitive cost of $2 to $6 per million tokens.
💡Why it mattersGrok 4.20, although less performant, may appeal due to its increased factual reliability, a major asset for applications requiring high precision.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Grok 4.20, the latest artificial intelligence model developed by xAI, struggles to compete with market leaders such as Gemini 3.1 Pro Preview and GPT-5.4. In benchmarks, Grok 4.20 scored 48 on the Intelligence Index, according to the report from Artificial Analysis. This figure is significantly lower than the 57 points achieved by its competitors, although Grok 4.20 improved its score by 6 points compared to its predecessor, Grok 4.

Despite its overall lower performance, Grok 4.20 stands out for its factual reliability. In the AA Omniscience test, which assesses a model's ability to avoid hallucinations and recall accurate facts, Grok 4.20 recorded a non-hallucination rate of 78%, a record according to Artificial Analysis. This means that the model was wrong only once in five times when it did not know the answer, a notable result in the field of AI.

xAI has introduced three API variants for Grok 4.20: one with reasoning, one without reasoning, and one in multi-agent mode. The model supports a context window of 2 million tokens, which is a significant advancement. In terms of cost, Grok 4.20 is offered at a competitive rate of $2 to $6 per million tokens, making it more affordable than its predecessor and competitive with Western models.

In summary, while Grok 4.20 may not be at the forefront in terms of raw performance, its ability to provide reliable factual information could make it a strategic choice for applications requiring high precision.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.