⚡
Brief IA
›

SPADE and Hawkeye Automate AI Training and Optimization

🔬 Research·Tom Levy·

SPADE and Hawkeye Automate AI Training and Optimization

SPADE and Hawkeye Automate AI Training and Optimization
⚡
Key Takeaways
1Hawkeye shows a speed gain of 18.9× compared to expert Triton cores on emerging attention variants.
2SPADE has been evaluated with three Qwen3 models fine-tuned via GRPO across 400 deployments of 25 environments each.
3METR observes a significant increase in reported vulnerabilities in 2026 compared to 2025, including in NVD and OSV.
💡Why it matters — These tools illustrate the growing automation of complex tasks in AI, while METR's metrics demonstrate concrete accelerations in certain scientific fields.
⚡Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

In cybersecurity, METR observes a significant increase in reported vulnerabilities, while contributions to mathematics remain difficult to measure and AI research shows no clear acceleration. Meanwhile, university teams present SPADE to generate training environments and Hawkeye to write GPU kernels tailored to hardware. An AI researcher calls for caution regarding the implications of scientific success.

Hawkeye: Speed Gains and Concerns About Implications

The researchers behind Hawkeye announce a geometric speedup of 18.9× compared to Triton kernels written by experts, focusing on emerging attention variants. Hawkeye also enables the production of good kernels for recent and poorly documented hardware. More broadly, systems like Hawkeye aim to enhance the performance of AI systems on AI R&D tasks. Concurrently, Julian Togelius, an AI researcher, expresses concerns about the implications of successful AI research. He believes these advancements raise complex philosophical questions and considers the situation delicate.

METR: Contrasting Accelerations Across Fields

METR has examined cybersecurity, mathematics, and AI research. The organization attributes significant contributions to AI in cybersecurity, more limited contributions in mathematics, and difficult-to-measure impacts in AI research. In cybersecurity, METR notes a sharp increase in reported vulnerabilities in 2026 compared to 2025, both in projects like cURL, OpenSSL, Firefox, and Microsoft, as well as in aggregated databases such as the NVD and OSV. In mathematics, METR observes an increase in the volume of work, with submissions on arXiv doubling in some fields in less than 12 months, while highlighting the difficulty in assessing the value of these contributions. Problems from recognized lists, including the Jacobian conjecture from Smale's list, problem 44 from Green's list (halving sieve), and the sofic half of problem 100 from Green, have been solved, but it is considered too early to determine if this momentum will be sustained. For AI research, METR analyzed seven algorithmic milestones, including CIFAR-10, Hutter compression, Gurobi MIP, MIPLIB, nanoGPT, Stockfish, and the matrix multiplication exponent. Contributions from LLMs are noted for nanoGPT and CIFAR-10, but the acceleration of usage remains much lower than in the other two fields.

SPADE: Environment Generation and Large-Scale Evaluation

Teams from several universities have developed SPADE, described as a versatile platform for creating co-evolutionary environments and agent skill acquisition through self-play. SPADE produces game-like environments, generating synthetic data to train LLMs by alternating between creating executable environments and solving them with a model. Two roles are assigned: a designer responsible for writing complete, long-duration environments in executable code, and a reasoning agent that learns to intervene. The reward for this agent is estimated by the performance gap with and without privileged hints. SPADE operates at a scale of 30B. Three Qwen3 models have been trained to test SPADE, each fine-tuned via GRPO across 400 deployments of 25 environments, and then evaluated on various benchmarks.

Involved Teams and Technical Targets of SPADE and Hawkeye

The development of SPADE involved researchers from the University of Washington, Stanford, Northeastern, Carnegie Mellon, MIT, the National University of Singapore, the National University of Seoul, Stevens Institute of Technology, and the University of Chicago. Hawkeye was designed by researchers from Harvard, Stanford, Together AI, and Caltech. Hawkeye aims to make coding agents aware of hardware specifics with minimal expert intervention by introducing a taxonomy of unit tests that is both minimalist and comprehensive. This taxonomy helps manage computation during testing and generates hardware-appropriate kernels.

⚡

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.