⚡
Brief IA
›

SpaceXAI's Grok 4.6 Challenges OpenAI and Anthropic in AI

🛠️ AI Tools·Tom Levy·

SpaceXAI's Grok 4.6 Challenges OpenAI and Anthropic in AI

SpaceXAI's Grok 4.6 Challenges OpenAI and Anthropic in AI
⚡
Key Takeaways
1SpaceXAI has launched Grok 4.6, an AI model that it compares to industry leaders like OpenAI and Anthropic.
2The performance of Grok 4.6 is highlighted by SpaceXAI, but the data comes from the company itself.
3Caution is advised when evaluating these figures, which require independent critical analysis.
💡Why it matters — The arrival of Grok 4.6 could intensify competition in the AI sector, but the reliability of the data remains to be verified.
⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Grok 4.6 from SpaceXAI Challenges OpenAI and Anthropic in AI

SpaceXAI has unveiled Grok 4.6 and positions it on par with the best models from OpenAI and Anthropic, backed by data. These figures, selectively compiled and published by SpaceXAI itself, warrant caution.

Two months after acquiring Cursor and following the launch of its Grok Bot agent, SpaceXAI continues to advance in the realm of cutting-edge models. Elon Musk's company presented Grok 4.6 on August 12, a model designed for long-duration agents, coding, and office work. The accompanying benchmark table places it in the same category as GPT-5.6 Sol and Claude Fable 5. However, caution is advised, and the rest of this article explains why.

A Claimed Equality, Gaps in the Performance

On the composite index from Artificial Analysis (nine aggregated evaluations), Grok 4.6 reportedly scores 61 points. This allows it to claim a perfect tie with GPT-5.6 Sol at its maximum reasoning level, and a slight lag of one point behind Fable 5 Max. The leap is significant compared to Grok 4.5, which was credited with 56 points on the same index, and the model even takes the lead in certain professional work tests like GDPVal-AA.

However, the overall picture resembles less a victory than an admission into the club. On Terminal-Bench, a command-line evaluation, Grok 4.6 would peak at 26%, while its two rivals exceed 34%. This lag would also be confirmed on DeepSWE, a software engineering test.

Most importantly, SpaceXAI clarifies that the competitors' scores come from their technical sheets and public rankings, without the rival models being re-executed under identical conditions. A benchmark table published by a vendor is akin to a report card written by the student: the grades may not necessarily be wrong, but the choice of subjects displayed is theirs. The industry has already experienced this situation with the controversy surrounding the benchmarks of Llama 4, which became a textbook case.

The Real Benchmark is on the Pricing Grid

This launch illustrates less a race for intelligence than a commercial strategy. Grok 4.6 is priced at $2 per million tokens for input and $6 for output. In comparison, GPT-5.6 Sol lists at $5 and $30 on the same grid, while Fable 5 rises to $10 and $50. For subscribed users, this may seem trivial. However, for companies using these models on a much larger scale, the stakes are quite different.

The model is integrated directly into Cursor, in Grok Build, and via the API, with a quota doubled in the first week as a full-scale trial period. This maneuver extends the acquisition of Cursor and the joint development initiated this summer. SpaceXAI is not looking to convince benchmark juries but rather the developers who pay the API bill at the end of the month.

The story also invites patience. Grok 4 had topped the Artificial Analysis rankings in July 2025, before the hierarchy was reorganized with subsequent competitor releases. Its free availability in France had already illustrated this approach of volume over prestige. For French developers, the equation of the day rests on a model that would cost several times less than its direct rivals, for performance touted as comparable. It remains for independent evaluations to confirm (or not) the second half of the promise.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.