Brief IA

Pentagon: AI Trained on Classified Data, Challenges and Risks

🔬 Research·Tom Levy·

Pentagon: AI Trained on Classified Data, Challenges and Risks

Pentagon: AI Trained on Classified Data, Challenges and Risks
Key Takeaways
1The Pentagon is exploring training AI on classified data to improve the accuracy of models used in military contexts.
2Companies like OpenAI and xAI may soon train their models on sensitive information, despite potential security risks.
3Aalok Mehta highlights the dangers of leaks of classified information but notes that secure infrastructures already exist to mitigate these risks.
💡Why it mattersThis initiative could transform the use of AI in the military field while posing significant security challenges.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

The Pentagon and AI Training on Classified Data

The Pentagon is considering creating secure environments to allow companies specializing in generative artificial intelligence to train their models on classified data. This initiative, revealed by MIT Technology Review, could mark a turning point in the use of AI within the U.S. military.

Currently, AI models like Claude from Anthropic are already being used to address questions in classified contexts, including target analysis in Iran. However, the idea of training these models directly on classified data represents a significant advancement that carries unique security risks. This would involve integrating sensitive information, such as surveillance reports or battlefield assessments, into the models, thereby bringing AI companies closer to classified data.

Agreements with AI Giants

A U.S. defense official, speaking anonymously, indicated that training AI models on classified data could enhance their accuracy and effectiveness for certain specific tasks. This announcement comes as the Pentagon has already signed agreements with companies like OpenAI and xAI, founded by Elon Musk, to deploy their models in classified environments. The Department of Defense is pursuing a new agenda aimed at becoming an "AI-focused" combat force, particularly in the context of rising tensions with Iran. However, the Pentagon has yet to officially comment on these training plans.

The training of the models would take place in secure data centers accredited to host classified government projects. A copy of an AI model would then be associated with classified data. While the Department of Defense retains ownership of the data, personnel from AI companies with the necessary security clearances could exceptionally access it.

Preliminary Assessment on Unclassified Data

Before allowing this training on classified data, the Pentagon first wants to assess the accuracy and effectiveness of the models by training them on unclassified data, such as commercially available satellite images.

For a long time, the military has used computer vision models to identify objects in images and sequences captured by drones and aircraft. Federal agencies have also awarded contracts to companies to train AI models on this type of content. Additionally, AI companies are developing language models (LLMs) and chatbots tailored to government needs, such as Claude Gov from Anthropic, designed to operate in multiple languages and secure environments. However, the recent comments from the defense official represent the first indication that companies like OpenAI and xAI could train government-specific versions of their models directly on classified data.

Risks of Training on Classified Data

Aalok Mehta, director of the Wadhwani AI Center at the Center for Strategic and International Studies, and former AI policy lead at Google and OpenAI, warns about the risks associated with training on classified data. According to him, the main danger lies in the possibility that classified information could be leaked to anyone using the model. This could pose a problem if multiple military departments, with varying classification levels and information needs, shared the same AI model.

Mehta illustrates this risk with an example: a model with access to sensitive human intelligence, such as the name of an agent, could accidentally disclose this information to a part of the Department of Defense that is not authorized to access it. This would create a security risk for the concerned agent, a risk that is difficult to mitigate if a model is used by multiple groups within the military.

Existing Secure Infrastructures

However, Mehta emphasizes that it is less complex to protect this information from the outside world. "If you set this up correctly, the risk of this data being leaked on the Internet or sent back to OpenAI is very low," he states. The government already possesses part of the necessary infrastructure. For example, the security giant Palantir has won significant contracts to create a secure environment allowing officials to query AI models on classified topics without sending the information back to AI companies. Nevertheless, using these systems for training remains an unprecedented challenge.

AI as a Strategic Tool for the Pentagon

Encouraged by a memo from Secretary of Defense Pete Hegseth in January, the Pentagon is striving to integrate AI more deeply into its operations. This includes its use in combat, where generative AI could classify target lists and recommend which to strike first, as well as in more administrative roles, such as drafting contracts and reports.

Many tasks currently performed by human analysts could be delegated to advanced AI models requiring access to classified data, Mehta explains. This could include learning to identify subtle cues in an image, as an analyst would, or linking new information to historical context. Classified data could be extracted from a vast amount of texts, audio, images, and videos in many languages collected by intelligence agencies.

It remains difficult to determine which specific military tasks would require AI models to be trained on such data, Mehta warns. "Obviously, the Department of Defense has many incentives to keep this information confidential and does not want other countries to know exactly what capabilities we have in this area."

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.