Brief IA

Sequent and FrontierCode: Secure AI and Advanced Coding

🔬 Research·Tom Levy·

Sequent and FrontierCode: Secure AI and Advanced Coding

Sequent and FrontierCode: Secure AI and Advanced Coding
Key Takeaways
1Sequent, founded by the UK AI Security Institute and Timaeus, aims to improve the alignment of superintelligent AIs.
2FrontierCode, created by Cognition, offers a coding benchmark with a score of 13.4% for Claude Opus 4.8.
3ChinaHeritaQA assesses the cultural capabilities of AI models with 2,279 images of Chinese heritage sites.
💡Why it mattersThese initiatives enhance the security and evaluation of AIs, which are crucial for their future integration into complex tasks.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Sequent: An Initiative for Aligning Superintelligent AIs

Researchers from the UK AI Security Institute and the startup Timaeus have launched a new nonprofit organization, Sequent, dedicated to researching the alignment of superintelligent artificial intelligence systems. Their mission is to develop techniques that enhance trust in the safety of these systems. According to them, while superintelligent AI may be developed in the coming years, it is uncertain whether alignment will keep pace. They emphasize that current programs from AI labs do not provide sufficient guarantees before training these systems.

Sequent plans to recruit between 40 and 80 full-time employees in the coming years and aims to raise between $100 and $150 million initially. They are also preparing to secure additional funding if their parallel research proves fruitful.

A Portfolio of Diverse Research

Sequent takes a distinct approach compared to large AI labs, seeking to establish sound reasons to believe that alignment observed in controlled environments can be generalized to more complex situations. Research areas include evolutionary monitoring, learning theory, heuristic arguments, game theory, and personas. The goal is to foster promising interactions between these different approaches.

The Importance of Alignment Before Autonomous Improvement

Currently, AI systems exhibit alignment flaws that can lead to unexpected failures. While these failures are acceptable at this stage, it becomes crucial to develop better alignment techniques as AI systems become more intelligent and humans begin to delegate more fundamental research to them.

ChinaHeritaQA: A Benchmark for Cultural Reasoning

Researchers from several universities have developed ChinaHeritaQA, a multimodal dataset designed to assess the cultural reasoning capabilities of vision-language models. This dataset includes 2,279 images from 51 UNESCO World Heritage sites in China, accompanied by 14,133 question-answer pairs in both Chinese and English. The questions cover various aspects such as identity recognition, visual anchoring, historical contextualization, and architectural analysis. Open-weight models already outperform humans, with an average accuracy of 67% for humans compared to 81% for the top-rated model.

FrontierCode: An Advanced Coding Benchmark

Cognition, the creator of Devin, has introduced FrontierCode, a particularly demanding coding benchmark. This benchmark consists of 150 tasks divided into three levels of difficulty: Diamond, Main, and Extended. The score of 13.4% achieved by Claude Opus 4.8 on the most difficult component illustrates the complexity of this benchmark. The languages tested include Python, Go, TypeScript, JavaScript, Java, and C/C++.

The Importance of Rigorous Assessments

Challenging assessments like FrontierCode are essential for measuring the rapid progress of AI. It is expected that systems will reach 70% or more on the Diamond level by June 2027. These benchmarks could last longer than other recent ones, providing a solid foundation for evaluating the capabilities of AI systems.

Xiaomi and Its Ultra-Fast Model

The technology company Xiaomi has unveiled the Xiaomi MiMo-V2.5-Pro-UltraSpeed, a model with 1 trillion parameters capable of processing 1,000 tokens per second. This performance is the result of co-designing the model with the software stack, utilizing techniques such as FP4 quantization.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.