Brief IA

GM Revolutionizes High-Speed Autonomous Driving AI

🤖 Models & LLM·Tom Levy·

GM Revolutionizes High-Speed Autonomous Driving AI

GM Revolutionizes High-Speed Autonomous Driving AI
Key Takeaways
1General Motors is developing an autonomous driving AI capable of processing scenarios at 50,000 times real speed.
2Vision Language Action models enable vehicles to understand complex situations, such as a police officer directing traffic.
3GM's On Policy Distillation technique significantly reduces the learning time for autonomous systems.
💡Why it mattersThis technological advancement could revolutionize the safety and efficiency of autonomous vehicles, preparing them for unforeseen and complex situations.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

An Approach to Evolving Autonomy

Autonomous driving is one of the most complex challenges for artificial intelligence. An automated system must interpret a chaotic and constantly evolving world in real-time. This involves navigating uncertainty, predicting human behavior, and operating safely in a variety of environments and edge cases.

At General Motors, the approach is based on a simple principle: while most moments on the road are predictable, it is the rare, ambiguous, and unexpected events — the long tail — that determine whether an autonomous system is safe, reliable, and ready for large-scale deployment.

Testing Long-Term Scenarios

Long-term autonomous driving scenarios come in various forms. Some are notable for their rarity. For example, a mattress on the road, a fire hydrant bursting, or a massive power outage in San Francisco that disabled traffic lights, forcing driverless vehicles to navigate challenges never encountered before. These rare system-level interactions, particularly in dense urban environments, demonstrate how unexpected edge cases can propagate on a large scale.

Training AI at Unprecedented Speeds

To accelerate the development of autonomous driving AI, GM has implemented a system that allows the AI to train at 50,000 times real speed. This revolutionary method enables the simulation of complex driving scenarios in record time, providing an unprecedented learning capability.

Deploying Visual Language Models

To address these nuanced scenarios, GM is developing Vision Language Action (VLA) models. Starting from a standard visual language model that leverages internet-scale knowledge to interpret images, GM engineers use specialized decoding heads to refine specific driving-related tasks. These models enable a vehicle to recognize that a hand gesture from a police officer takes precedence over a red light or to identify what a "loading zone" looks like in a busy airport terminal.

Testing Dangerous Scenarios in High-Fidelity Simulations

Driving requires reaction times in a fraction of a second, so any excessive latency poses a critical problem. To address this, GM is developing a dual-frequency VLA. This large-scale model operates at a lower frequency to make high-level semantic decisions, while a smaller, highly efficient model manages immediate spatial control. This hybrid approach allows the vehicle to benefit from deep semantic reasoning without sacrificing the reaction times necessary for safe driving.

Synthetic Data for the Toughest Cases

The simulated scenarios come from various AI technologies used by GM engineers to produce innovative training data. For example, GM's "Seed-to-Seed Translation" research leverages diffusion models to transform existing real data, allowing clear recordings to be modified into rainy or foggy scenes while preserving the scene's geometry.

From Abstract Policy to Real-World Driving

To bring this conceptual expertise into the real world, GM is one of the first to use a technique called On Policy Distillation. Engineers run their simulator in simultaneous modes: the abstract, high-speed mode, and the high-fidelity sensor mode. Here, the reinforcement learning model, which has practiced countless abstract kilometers to develop a perfect "policy," acts as a teacher. This knowledge transfer is incredibly efficient; just 30 minutes of distillation can capture the equivalent of 12 hours of raw reinforcement learning.

Designing Failures Before They Occur

The simulation aims not only to train the model to drive well but also to make it fail. To rigorously test the system, GM uses a differentiable pipeline called SHIFT3D. This pipeline actively modifies the world to create "adversarial" objects designed to deceive the perception system.

Conclusion

This rigorous and multifaceted approach — from the "Boxworld" strategy to adversarial stress testing — constitutes the framework proposed by General Motors to solve the final 1% of autonomy. While it serves as a foundation for future development, it also raises new research challenges that engineers must tackle.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.