In a Manhattan courtroom, a high-stakes legal drama is unfolding, pitting the titans of artificial intelligence against the stalwart institutions of journalism.
At its heart is a battle not just over copyright, but over the very nature of truth, compensation, and accountability in an increasingly AI-driven world.
The New York Daily News, The New York Times, and a consortium of other prominent news outlets are demanding that ChatGPT’s parent company, OpenAI, cease its alleged practice of systematically deleting data.
They contend this data is crucial evidence of widespread intellectual property theft.
The accusations leveled against OpenAI are stark and unyielding.
Lawyers for the news organizations assert that the tech behemoth has employed “every trick in the book” to obscure its tracks.
These range from mass data deletions to the implementation of filters ostensibly designed to make it harder for OpenAI’s products to “regurgitate copyrighted content.” This isn’t just a technical dispute; it’s an allegation of digital larceny on an unprecedented scale.
Such actions are threatening to undermine the economic foundations of investigative journalism.
Central to the current skirmish is a directive from Manhattan Federal Magistrate Judge Ona Wang.
Last month, she ordered OpenAI to preserve its output logs and related information that were slated for deletion.
This order was a direct response to the news outlets’ claims that OpenAI was deliberately scrubbing vast swaths of data.
This scrubbing allegedly hindered their ability to demonstrate how AI products could circumvent paywalls and plagiarize their meticulously reported work.
OpenAI, however, has pushed back, appealing to Judge Wang to vacate her order.
Their argument? That continuing to store this data would constitute a “massive burden” and, somewhat ironically, infringe upon the privacy of its users.
For the news outlets, this privacy defense rings hollow, bordering on the disingenuous.
They highlight that OpenAI’s own user agreements often stipulate data retention when legally required, making their current stance appear contradictory.
Moreover, they point out that OpenAI has not denied the relevance of the deleted data to the ongoing lawsuits.
“What it does not dispute is that the output log data is relevant to the News Cases,” the lawyers declared.
They added a pointed observation: that a company valued at more than $300 billion, a global leader in technological innovation, undeniably possesses “both the means and ability to preserve this concededly relevant data.” The implication is clear: OpenAI’s plea of “burden” is a flimsy veil for a more sinister motive – to destroy evidence.
The core of the news outlets’ complaint goes far beyond mere data retention.
They allege that OpenAI has illegally harvested millions of news stories.
These stories are the product of billions of dollars spent by publishers sending “real people to real places to report on real events in the real world.” This vast reservoir of human endeavor, they argue, has been “purloined” without compensation to train OpenAI’s large language models.
This results in generative AI products that can “vomit out” – or misrepresent – pirated reporting to users.
The consequences are profound, not just for the economic viability of newsrooms, but for the very integrity of information.
When AI misstates or misrepresents original reporting, it risks misinforming ChatGPT users and eroding public trust in both the source and the medium.
OpenAI, for its part, has sought refuge under the umbrella of fair use.
This is a legal doctrine that permits the limited use of copyrighted material for purposes such as criticism, commentary, news reporting, teaching, or research.
However, the news outlets fiercely contest this defense.
They contend that the fair use test typically requires a copyrighted work to be transformed into something new.
Crucially, they argue that the new work cannot compete with the original in the same marketplace.
OpenAI’s AI products, they argue, do precisely that: they provide news and information.
This often directly competes with the very sources they have allegedly stolen from, without any compensatory mechanism.
The legal skirmish has also seen OpenAI attempt to dismiss the notion that its products are being used as a substitute for traditional news consumption.
They had previously argued that the newspapers hadn’t produced “a shred of evidence” to support this claim.
Yet, the news outlets have swiftly countered, noting that OpenAI’s own engineers have all but admitted as much.
They cited an instance in another lawsuit involving Google where an OpenAI engineer acknowledged that “local news” was a “pretty common query” among ChatGPT users.
This effectively conceded that their chatbots, even if not explicitly designed to slip past paywalls, were indeed functioning as a source of news for many.
This legal battle, initiated by The New York Times in December 2023 and joined by the Daily News and other MediaNews Group and Tribune Publishing outlets in April 2024, is more than a simple copyright dispute.
It is a defining moment for the future of intellectual property in the age of artificial intelligence.
It forces a reckoning with how technology, no matter how transformative, must adhere to established ethical and legal frameworks.
The outcome of Judge Wang’s decision on data preservation, and ultimately the broader lawsuit, will send a powerful message.
This message is about whether the digital realm will truly be a wild west, or if the foundational principles of creative endeavor and fair compensation will endure.
As news organizations fight for their very survival in a rapidly evolving digital landscape, the stakes in this courtroom drama could not be higher.
-
Frank DiBernardo handles LNGFRM's Foodie and Miscellaneous writing tasks. He's always getting ideas from users, so don't be afraid to send an email to the editor.