Microsoft Exploits Unlicensed Web Data for Its AI
Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
Microsoft and the Use of Unlicensed Data
Microsoft has recently come under scrutiny for training its AI models partially on unlicensed web data. A technical document, reviewed by Simon Willison, reveals that Microsoft has utilized sources such as Common Crawl to feed its models, despite previous statements claiming that the data used was "clean, enterprise-grade, and commercially licensed."
The Use of Fair Use
Like many other companies in the artificial intelligence field, Microsoft appears to rely on the concept of fair use to justify the use of this data. The technical document describes the data as a mix of human-generated content, publicly accessible, and licensed. To collect this data, Microsoft employs a proprietary crawler that adheres to the Robots Exclusion Protocol (robots.txt) as well as associated HTML meta-tags and controls. This allows website owners to decide how their content is accessed and used.
The Responsibility of Website Owners
This approach places the responsibility for content protection on website owners. In other words, those who do not take measures to protect their content could be considered to consent to its use. However, the application of fair use in this context remains a subject of debate, and courts have not yet definitively ruled on this issue.
A Common Yet Contested Practice
Ultimately, Microsoft seems to be following a common practice among AI companies while presenting its training data as particularly "clean." This situation raises questions about the transparency and ethics of data collection practices by tech giants.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.