Anthropic fined $1.5 billion, training deemed legal

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
The framework of copyright law surrounding AI is being built gradually. A court has validated the legality of training Anthropic's models but sanctioned the use of pirated sources, while another court rejected fair use when the goal was to compete with a major information services company. Furthermore, a work that is 100% generated by AI is not eligible for protection. Between these milestones and numerous ongoing procedures, no definitive answers are forthcoming at this time.
Ongoing Disputes and Provisional Decisions
Most AI companies are currently involved in legal proceedings on these issues, and no definitive solution is expected in the short term. Cathy Gellis believes that the initial decisions carry weight, but they could be contradicted if other jurisdictions rule differently. She anticipates that later stages of litigation will determine which position will prevail. According to her, these ongoing decisions are already shaping the situation, and it would be unwise for AI companies to ignore them.
AI-Generated Works: Protection Denied if Creation is 100%
Cathy Gellis distinguishes the debate over model training from that concerning the content produced by AI. In Thaler v. Perlmutter, the court ruled that a work that is 100% generated by AI cannot be protected by copyright. This ruling raises practical difficulties: establishing that a work was generated with AI and, if applicable, determining what percentage falls under creation or assistance by AI. Gellis reminds us that using a spell checker in Microsoft Word does not grant ownership of the text to the software publisher, and she believes that AI necessitates revisiting many previously accepted decisions.
Fair Use: Legal Criteria and Competitive Weight
Many disputes revolve around "fair use," meaning the use of protected content without permission when it is sufficiently transformative. Courts particularly assess the purpose and nature of the use, the amount taken, and the impact on the market. Jason Henderson notes that copyright aims to protect and grow the market, while observing varied reasoning among courts. According to him, when training is used to directly compete, jurisdictions tend to be unfavorable, and when there is no competition, they tend to find the use acceptable.
Ross Case: No Fair Use for Building a Competitor
Thomson Reuters sued Ross Intelligence for copying content to create a competing AI-based legal platform. Judge Stephanos Bibas determined that Ross's use was not transformative and concluded that there was no fair use in this scenario of direct competition. The argument, sometimes put forward, that chatbots would compete with authors by generating synthetic books from their works has, at this stage, not convinced the courts.
Anthropic: Training Found Legal, Illicit Sources Penalized
In one of the first rulings of its kind, Judge William Alsup ordered Anthropic to pay $1.5 billion to writers whose works had been used, while deeming the training of its models legal. The penalty targeted the illicit acquisition of books taken from illegal online libraries. Alsup described models trained to create something other than copies and compared the ingestion of trillions of words to an author studying literature. For Cathy Gellis, this framework is rather favorable to AI companies: she puts into perspective a $1.5 billion fine against a company projected to generate around $200 billion in annual revenue by 2028 and is pleased that the judge equated training with reading rather than copying. She emphasizes that, in her view, copyright is based on copying, not on the use or reading of a work.
A 1976 Law Facing Massive Training Corpora
The U.S. copyright law has not been revised since 1976, meaning judges are navigating with guidelines that are about 50 years old in the face of recent AI uses. The large models that power ChatGPT, Gemini, or Claude are trained on vast corpora: hundreds of millions of books, online articles, academic publications, and, more broadly, content available on the Internet. Many published authors may have contributed without being informed or consenting. For Cathy Gellis, a practitioner of copyright and technology law, this context fuels a complex debate with opposing sensitivities.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.