Sequent and FrontierCode: Secure AI and Advanced Coding
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Sequent: An Initiative for Aligning Superintelligent AIs
Researchers from the UK AI Security Institute and the startup Timaeus have launched a new nonprofit organization, Sequent, dedicated to researching the alignment of superintelligent artificial intelligence systems. Their mission is to develop techniques that enhance trust in the safety of these systems. According to them, while superintelligent AI may be developed in the coming years, it is uncertain whether alignment will keep pace. They emphasize that current programs from AI labs do not provide sufficient guarantees before training these systems.
Sequent plans to recruit between 40 and 80 full-time employees in the coming years and aims to raise between $100 and $150 million initially. They are also preparing to secure additional funding if their parallel research proves fruitful.
A Portfolio of Diverse Research
Sequent takes a distinct approach compared to large AI labs, seeking to establish sound reasons to believe that alignment observed in controlled environments can be generalized to more complex situations. Research areas include evolutionary monitoring, learning theory, heuristic arguments, game theory, and personas. The goal is to foster promising interactions between these different approaches.
The Importance of Alignment Before Autonomous Improvement
Currently, AI systems exhibit alignment flaws that can lead to unexpected failures. While these failures are acceptable at this stage, it becomes crucial to develop better alignment techniques as AI systems become more intelligent and humans begin to delegate more fundamental research to them.
ChinaHeritaQA: A Benchmark for Cultural Reasoning
Researchers from several universities have developed ChinaHeritaQA, a multimodal dataset designed to assess the cultural reasoning capabilities of vision-language models. This dataset includes 2,279 images from 51 UNESCO World Heritage sites in China, accompanied by 14,133 question-answer pairs in both Chinese and English. The questions cover various aspects such as identity recognition, visual anchoring, historical contextualization, and architectural analysis. Open-weight models already outperform humans, with an average accuracy of 67% for humans compared to 81% for the top-rated model.
FrontierCode: An Advanced Coding Benchmark
Cognition, the creator of Devin, has introduced FrontierCode, a particularly demanding coding benchmark. This benchmark consists of 150 tasks divided into three levels of difficulty: Diamond, Main, and Extended. The score of 13.4% achieved by Claude Opus 4.8 on the most difficult component illustrates the complexity of this benchmark. The languages tested include Python, Go, TypeScript, JavaScript, Java, and C/C++.
The Importance of Rigorous Assessments
Challenging assessments like FrontierCode are essential for measuring the rapid progress of AI. It is expected that systems will reach 70% or more on the Diamond level by June 2027. These benchmarks could last longer than other recent ones, providing a solid foundation for evaluating the capabilities of AI systems.
Xiaomi and Its Ultra-Fast Model
The technology company Xiaomi has unveiled the Xiaomi MiMo-V2.5-Pro-UltraSpeed, a model with 1 trillion parameters capable of processing 1,000 tokens per second. This performance is the result of co-designing the model with the software stack, utilizing techniques such as FP4 quantization.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.