NEWS

News Outlets Demand OpenAI Data Preservation

Leading news outlets demand OpenAI halt data deletions, claiming the AI giant is destroying crucial evidence of alleged copyright infringement. Publishers argue OpenAI’s actions are a deliberate attempt to conceal widespread plagiarism of their content.

By
LNGFRM Team
Published June 19, 2025
Stylized illustration of a tablet displaying a circular scanner icon with a green center and a blue padlock, set against winding grey paths on a green background.
Illustration by Addison Smith for LNGFRM

In the hallowed halls of Manhattan’s federal courthouse, a battle of unprecedented significance is unfolding, one that pits the bedrock of traditional journalism against the burgeoning titan of artificial intelligence.

At its heart lies a fundamental question: can a technological marvel built on the world’s collected knowledge evade accountability for the very content it consumes and repurposes?

This week, a coalition of America’s leading news organizations, including The Daily News and The New York Times, amplified their urgent plea to Manhattan Federal Magistrate Judge Ona Wang.

Their demand is clear and resonant: reject OpenAI’s audacious attempt to continue deleting vast swaths of data, information that they argue is crucial evidence of the AI giant’s alleged intellectual property theft.

It’s a move the news outlets describe as nothing less than using “every trick in the book” to conceal what they contend is pervasive plagiarism.

The current skirmish centers on Judge Wang’s order last month, which mandated OpenAI to preserve its output logs and related data.

This directive came in response to accusations that ChatGPT’s parent company was systematically purging enormous quantities of information, effectively hamstringing the news outlets’ ability to demonstrate how OpenAI’s products allegedly circumvent paywalls to “plagiarize and regurgitate copyrighted content.”

OpenAI, predictably, has pushed back, imploring Judge Wang to vacate her order.

Their rationale? The continuous storage of such data would impose a “massive burden” and, rather conveniently, infringe upon user privacy.

Yet, the news outlets are quick to expose the apparent inconsistencies in OpenAI’s defense.

They point out that the company’s own terms of service often stipulate data retention when legally required.

More damningly, they highlight that OpenAI has not, for a moment, denied the relevance of the data being deleted to the ongoing lawsuit.

“What it does not dispute is that the output log data is relevant to the News Cases, which as OpenAI has long recognized, include infringement claims based on outputs generated by [OpenAI’s] models and products,” lawyers for the publishers asserted in court documents.

The irony is not lost on observers: a company now valued at over $300 billion, a technological juggernaut with seemingly limitless resources, claims a “massive burden” in preserving data directly pertinent to a multi-billion-dollar legal challenge.

This isn’t merely a dispute over data storage; it’s an accusation of deliberate obfuscation.

The news outlets contend that the mass deletions are part of a broader strategy to “skirt accountability.”

They’ve also accused OpenAI of installing digital filters “designed to make it harder” to elicit answers containing copyrighted journalistic works.

This, they argue, is an implicit admission of guilt—why install filters if there’s no problematic content to filter?

“OpenAI’s preferred course of action to ‘protect its users’ data and privacy’ — immediately resuming mass deletions — will also, coincidentally, allow it to continue to destroy data that would show its liability for copyright infringement,” the news outlets’ lawyers wrote, laying bare the perceived motive.

Judge Wang, in her May 13 order, had already pre-emptively addressed OpenAI’s privacy concerns, clarifying that the data would be preserved and segregated, not provided “wholesale” to anyone or stored “forever.”

It was a measured approach designed solely to address the concerns raised in the suit.

Should the judge, however, be inclined to entertain OpenAI’s objection further, the newspapers have requested an opportunity to analyze different populations of data and present their findings to the court, a testament to their determination to unearth the truth.

The core of the lawsuit itself paints a stark picture of alleged digital piracy on an industrial scale.

The plaintiffs accuse OpenAI of illegally harvesting millions of news stories, the very lifeblood of their operations, to train its large language models.

These models, they argue, then build generative AI products that can “vomit them out—or versions of them—to users.”

The consequences, the newspapers contend, are not merely financial.

They include the disturbing phenomenon of journalists’ pirated reporting being misstated or misrepresented, thereby actively misinforming ChatGPT users.

This legal showdown cuts to the very heart of journalistic enterprise.

Publishers, the lawsuit eloquently notes, have collectively spent “billions of dollars to send real people to real places to report on real events in the real world.”

OpenAI, in stark contrast, is accused of “purloining” this invaluable reporting without compensation, solely “to create products that provide news and information plagiarized and stolen.”

OpenAI’s primary defense rests on the doctrine of fair use, a legal concept that permits the limited use of copyrighted material for purposes such as criticism, commentary, news reporting, teaching, and research.

However, the news outlets vigorously dispute this application.

They argue that the fair use test mandates a transformation of the copyrighted work into something genuinely new, and crucially, the new work cannot compete with the original in the same marketplace.

By generating news summaries or information derived directly from their articles, the plaintiffs contend, OpenAI’s products directly compete with, and undermine, their own offerings.

Further bolstering their case, the news outlets highlighted how Judge Wang had previously rejected OpenAI’s assertion that they hadn’t produced “a shred of evidence” that people are using ChatGPT or OpenAI’s API products to get news instead of paying for it.

In a telling moment, the newspapers pointed out that OpenAI engineers had all but admitted as much, acknowledging that while their chatbots weren’t designed to slip past paywalls, they certainly could.

They also cited another ongoing suit involving Google, where an OpenAI engineer conceded that local news was a “pretty common quer[y]” among ChatGPT users.

Such admissions underscore the practical reality that AI models, irrespective of their stated intent, are indeed becoming a source of news for many, often at the expense of the original creators.

This complex legal saga, initiated by The New York Times in December 2023, and later joined by The Daily News and other prominent outlets like The Mercury News, The Denver Post, and the Chicago Tribune in April 2024, represents more than just a copyright dispute.

It is a defining moment for the future of intellectual property in the digital age, a test of whether AI innovation can truly flourish without respecting the rights and livelihoods of those who create the foundational content it consumes.

As the legal maneuvering continues, with OpenAI’s lawyers notably silent on The News’ requests for comment, the world watches to see if the titans of tech will be held accountable for the very information that fuels their meteoric rise.

Author

  • LNGFRM Team

    Frank DiBernardo handles LNGFRM's Foodie and Miscellaneous writing tasks. He's always getting ideas from users, so don't be afraid to send an email to the editor.

Daily Newsletter
Subscribe to our Newletter!
You May Also Like
© 2026 LNGFRM. All rights reserved.