Brief IA

Microsoft Exploits Unlicensed Web Data for Its AI

🤖 Models & LLM·Tom Levy·

Microsoft Exploits Unlicensed Web Data for Its AI

Microsoft Exploits Unlicensed Web Data for Its AI
Key Takeaways
1Microsoft used Common Crawl data to train its MAI models, despite promises of licensed data.
2The technical document reveals that Microsoft relies on fair use to justify the use of this data.
3Site owners must use protocols like robots.txt to protect their content from Microsoft's crawlers.
💡Why it mattersThis practice raises questions about the transparency and ethics of tech giants in their use of web data.
Le brief IA que lisent les pros

Le brief IA que les pros lisent chaque soir

Les 7 actus IA du jour, décryptées en 5 min. Gratuit.

Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.

Choisis ton rythme

Gratuit · Pas de spam · Désabonnement en 1 clic

📄
Full Analysis

Microsoft and the Use of Unlicensed Data

Microsoft has recently come under scrutiny for training its AI models partially on unlicensed web data. A technical document, reviewed by Simon Willison, reveals that Microsoft has utilized sources such as Common Crawl to feed its models, despite previous statements claiming that the data used was "clean, enterprise-grade, and commercially licensed."

The Use of Fair Use

Like many other companies in the artificial intelligence field, Microsoft appears to rely on the concept of fair use to justify the use of this data. The technical document describes the data as a mix of human-generated content, publicly accessible, and licensed. To collect this data, Microsoft employs a proprietary crawler that adheres to the Robots Exclusion Protocol (robots.txt) as well as associated HTML meta-tags and controls. This allows website owners to decide how their content is accessed and used.

The Responsibility of Website Owners

This approach places the responsibility for content protection on website owners. In other words, those who do not take measures to protect their content could be considered to consent to its use. However, the application of fair use in this context remains a subject of debate, and courts have not yet definitively ruled on this issue.

A Common Yet Contested Practice

Ultimately, Microsoft seems to be following a common practice among AI companies while presenting its training data as particularly "clean." This situation raises questions about the transparency and ethics of data collection practices by tech giants.

Brief IA — L'actualité IA en français

L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.