OpenAI Accused of Hiding Evidence by the New York Times

Le brief IA que les pros lisent chaque soir
Les 7 actus IA du jour, décryptées en 5 min. Gratuit.
Inclus dès l'inscription : notre sélection des meilleurs guides & comparatifs IA.
Choisis ton rythme
Gratuit · Pas de spam · Désabonnement en 1 clic
OpenAI Under Fire from the New York Times
The New York Times, supported by The Daily News, accuses OpenAI of deliberately hiding its ability to analyze user data and training datasets to detect copyrighted works. This accusation is part of a two-year-long lawsuit, where OpenAI is suspected of using protected content from the Times to train its generative AI models, including ChatGPT, and reproducing this content in the responses provided to users.
OpenAI's Arguments and Media Requests
Throughout this legal proceeding, OpenAI has claimed that it is impossible for the company to search its own training corpus due to technical complexity and user privacy concerns. The company explained that accessing this data would require retrieving, processing, and anonymizing it, which would pose a considerable logistical challenge. Meanwhile, the media has requested access to this data to verify whether their protected content was used by OpenAI and how frequently ChatGPT reproduced their works.
Revelations During a Deposition
During a deposition in April, Vinnie Monaco, a data protection engineer at OpenAI, reportedly revealed that the company had already conducted internal research on its training corpus to find copyrighted works. This statement contradicts OpenAI's previous claims about its inability to analyze this data.
A Secret Database and a Detection Filter
Monaco's deposition also revealed that even before the NYT filed its complaint, OpenAI had created a database containing approximately 78 million anonymized ChatGPT conversations. This database was used to assess the extent of potential copyright infringements. Additionally, OpenAI reportedly implemented a filter named "Bloom" as part of a project called "Project Giraffe," aimed at detecting and recording repetitions in the responses generated by ChatGPT.
The Implications of the Revelations
These revelations are crucial as they show that OpenAI had the capability to provide information that the plaintiffs had requested from the outset of the case. Initially, the plaintiffs had requested a sample of 120 million chat logs, but OpenAI managed to reduce this sample to just 20 million. However, this sample, submitted last December, was so heavily redacted that it became unusable, according to the court.
Accusations of Evidence Destruction
The plaintiffs also accuse OpenAI of deleting billions of ChatGPT responses after the complaint was filed, which would constitute a violation of the court's preservation order. They claim that OpenAI replaced millions of logs in the requested sample, making access to the information unnecessarily complex.
Reactions and Requests from the Plaintiffs
Ian B. Crosby, lead attorney for the plaintiffs, stated that if OpenAI truly believed that the use of their clients' journalism was legal, it would not have hidden the truth. The NYT and The Daily News are now asking the judge to sanction OpenAI for disrupting the discovery process and preventing the use of the 20 million log sample as evidence, deeming it unreliable. They also request that the court accept as fact that the ChatGPT logs would have shown significant repetition and a basis of the plaintiffs' content, and that it prevent OpenAI from arguing that its provided chat logs do not demonstrate substantial repetition. Finally, they seek to have OpenAI held responsible for legal fees incurred in tracking this evidence.
OpenAI's Response
Drew Pusateri, a spokesperson for OpenAI, dismissed these accusations, stating that the Times is seeking access to private user conversations as its case weakens. He claimed that the Times persists in its efforts to violate user privacy by making allegations he describes as patently false. OpenAI asserts that it continues to defend user privacy and the principles of fair use.
Brief IA — L'actualité IA en français
L'essentiel de l'actualité de l'intelligence artificielle, décrypté et expliqué chaque jour.