Brief IA

SpaceXAI, OpenAI, Meta: The Race for AI Models Heats Up

🔬 Research·Tom Levy·

SpaceXAI, OpenAI, Meta: The Race for AI Models Heats Up

SpaceXAI, OpenAI, Meta: The Race for AI Models Heats Up
Key Takeaways
1SpaceXAI, OpenAI, and Meta have each launched new AI models, marking a week of significant developments in the sector.
2Closed models, such as Grok 4.5 and GPT-5.6, now outperform open-weight models in terms of cost per unit of intelligence.
3OpenAI reported a rapid increase in active users, reaching 8 million thanks to the integration of Codex and ChatGPT Work.
💡Why it mattersThis heightened competition among AI giants could transform the economics of AI models, influencing costs and accessibility for users.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Launch of New AI Models by SpaceXAI, OpenAI, and Meta

This week has been particularly dynamic in the field of artificial intelligence with the launch of several major models by key players in the industry. SpaceXAI unveiled its model Grok 4.5, while OpenAI expanded access to GPT-5.6, moving from a restricted preview to general availability. Meanwhile, Meta introduced Muse Spark 1.1. OpenAI also presented GPT-Realtime-2.1 along with a mini version, which enhance interruption management and alphanumeric recognition, while reducing the p95 latency by at least 25%.

Three new leading competitors have emerged almost simultaneously, marking a turning point in the competition. The most notable development of the week concerns a recalibration of pricing and performance. Closed models, which already had an advantage in terms of raw intelligence, are seeing their cost advantage per unit of intelligence diminish. GLM-5.2, a leading open-weight model, scores 51 at a cost of about $0.37 per benchmark task. In comparison, Grok 4.5 achieves a score of 54 at a cost of $0.31, surpassing GLM in both score and cost. GPT-5.6 and Muse Spark 1.1 match GLM with a score of 51, but cost approximately $0.21 and $0.26 per task, respectively. The models Sol and Terra achieve scores of 59 and 55, but at a higher task cost. Closed models now match or exceed open-weight models in terms of cost per unit of intelligence, a major selling point for the latter.

Comparison of Costs and Performance of AI Models

Cost per task has become a more relevant indicator than price per million tokens. For example, Grok 4.5 charges more per output token than GLM-5.2, but uses about 14,000 output tokens per benchmark task compared to 43,000 for GLM. This means that Grok's intelligence per token compensates for a higher token price. An identical SVG prompt from Simon Willison cost 0.71 cents on Luna with reasoning disabled and 48.55 cents on Sol at maximum effort, revealing a 68x gap due solely to parameters.

State of Models and Their Adoption

Two weeks ago, we already discussed the preview of GPT-5.6, including the models Sol, Terra, and Luna, as well as pricing, security restrictions, and Ultra mode. The difficult-to-interpret result of METR was also mentioned. This week's news is more solid: the family is now generally available, and independent tests have confirmed the capabilities and efficiency claimed by OpenAI. Sol is once again in serious competition with Claude Fable 5. It scores 59 on the Artificial Analysis Intelligence Index compared to 60 for Fable, and leads the Coding Agents Index. Fable remains ahead on SWE-Bench Pro, GDPval, Toolathlon, FrontierMath Tier 4, and other professional comparisons. The choice between these models will depend on the task at hand. Sol excels in coding, computing tasks, presentations, and structured execution. In personal use, Fable is still preferred for writing, as shown by the preliminary ranking of Arena in creative writing, with a score of 1507 compared to 1486 for Sol, although uncertainty overlaps these results.

Luna could be the most useful launch of the family. It matches the frontier of open-weight intelligence, scoring 75 on the Coding Agents Index via Codex, operates at over 200 output tokens per second, and costs less per measured task than GLM-5.2. CodeRabbit warns about the other end of the family: Sol successfully completed 63.7% of over 100 submission tasks without execution errors, but its code review accuracy was only 31.6%. Long-term execution improves faster than verification.

The adoption of the new models has progressed almost as quickly as their development. OpenAI reported over 5 million weekly active users of Codex in early June. After the launch of ChatGPT Work and the merging of Chat, Work, and Codex into a single desktop application, the head of Codex reported 8 million active users across Codex and ChatGPT Work. Although definitions vary, this indicates a positive trend. Over a million people were already using Codex outside of software development before the launch of Work, and the new interface offers this audience more user-friendly access to the same agent infrastructure. This could be the strongest signal of the launch: the distribution layer is catching up with the capability layer.

Performance and Outlook of Grok 4.5

Grok 4.5 has closed the gap faster than expected. It now sits in the same broad group of coding agents as GPT-5.5 and Fable, while costing much less. It operates at about 80 tokens per second, achieves good scores on SWE Marathon and Terminal-Bench, and uses significantly fewer tokens than several peers. Snorkel measured a success rate of 29% across all criteria over approximately 2,000 professional tasks, ahead of GPT-5.5 and Opus 4.8. Grok operated within its own Grok Build framework, while competitors used a different agent, which must be taken into account as a system comparison. It still lags behind Fable and GPT-5.5 on DeepSWE 1.1. Cursor reported a contaminated result on CursorBench, and its hallucination rate of 54% on AA-Omniscience means it requires supervision. Grok is considered a credible frontier model with excellent economics, although a level below the very best in coding.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.