source: techcrunch ai: new york times says openai hid evidence in chatgpt copyright trial

level: business

the new york times and the daily news allege openai misled the court about its technical capabilities in their copyright lawsuit. the publishers say openai claimed it could not easily search its training data or chat logs, but a deposition revealed the company had already built internal tools to do so. an engineer testified that openai had a database of 78 million de-identified chatgpt conversations and a filter called bloom to detect regurgitated content.

the plaintiffs argue openai made discovery unnecessarily difficult. they had requested 120 million chat logs but accepted 20 million after negotiations. the sample openai provided was heavily redacted and deemed unusable by the court. the publishers also claim openai deleted billions of chatgpt outputs after the lawsuit was filed, violating a preservation order, and substituted millions of logs in the sample.

the times and daily news are now asking the judge to sanction openai. they want the court to bar openai from using the 20 million log sample as evidence, accept that chat logs would show major regurgitation, and prevent openai from arguing otherwise. they also seek legal fees. openai denies the allegations, calling them false and accusing the times of trying to invade user privacy as its case weakens.

why it matters: this case could set precedents for how ai companies handle discovery and transparency in copyright disputes, affecting data access for future litigation.


source: techcrunch ai: new york times says openai hid evidence in chatgpt copyright trial