NEWS

AI Training Moves to Library Archives

Artificial intelligence models are now training on vast library archives, moving beyond the internet’s unverified data. This shift, driven by legal battles and a quest for deeper knowledge, leverages centuries of human thought to create more accurate and responsible AI systems.

By
LNGFRM Team
Published June 16, 2025
Abstract illustration of data center servers connected by network lines to cloud symbols on a red background.
Illustration by Addison Smith for LNGFRM

The digital wild west, it turns out, was merely the prologue.

For years, artificial intelligence models gorged on the sprawling, often unverified, data of the internet – the fleeting chatter of social media feeds, the collaborative chaos of Wikipedia, and even vast, illicit troves of pirated books.

This diet, while abundant, left AI with a rather limited, and at times ethically dubious, understanding of humanity.

Now, in a profound shift, tech giants are turning to a more ancient, revered repository of knowledge: the hallowed shelves of libraries.

From the quiet sanctity of Harvard University to the meticulous archives of the Boston Public Library, centuries of human thought, meticulously preserved in ink and paper, are becoming the new frontier for AI training.

This pivot isn’t just about intellectual curiosity; it’s a pragmatic evolution, driven by a confluence of legal battles, data exhaustion, and a quest for deeper, more authentic understanding.

The legal landscape, in particular, has become a minefield for AI developers. Novelists, visual artists, and a growing chorus of creators have launched a barrage of lawsuits, alleging their copyrighted works were used without consent to train these powerful AI systems. The legal issues presented by generative AI are complicating this evolving landscape.

The solution, it seems, lies in the public domain. “It’s a prudent decision to start with public domain information because that’s less controversial at this moment than content that’s still copyrighted,” explained Burton Davis, Assistant General Counsel at Microsoft, a company heavily invested in AI development.

This move sidesteps immediate legal skirmishes, offering a safer, albeit more challenging, path to data acquisition.

Beyond legal expediency, there’s a deeper, more fundamental driver: the pursuit of quality and depth. The internet, for all its colossal size, is a relatively recent phenomenon, largely reflecting the last few decades of human discourse.

It lacks the cultural, historical, and linguistic richness embedded in centuries of published works. As Davis noted, libraries safeguard “enormous amounts of interesting cultural, historical, and linguistic data” that simply aren’t present in the contemporary online chatter AI has predominantly learned from.

Furthermore, the fear of data exhaustion is real, pushing developers to generate lower-quality “synthetic” data from the AI itself—a less-than-ideal solution that risks perpetuating errors and biases.

This paradigm shift is being facilitated by initiatives like Harvard’s Institutional Data Initiative, bolstered by “unrestricted gifts” from Microsoft and OpenAI, the company behind ChatGPT.

Their goal isn’t merely to extract data but to collaborate with libraries and museums globally, ensuring these historical collections are AI-ready in a way that also benefits the communities they serve.

“We’re trying to move some of the power that’s in AI’s hands right now back to these institutions,” stated Aristana Scourtas, who leads research at Harvard Law School’s Library Innovation Lab. “Librarians have always been the stewards of data and information.”

This partnership subtly rebalances the power dynamic, giving libraries a say in how their invaluable holdings are utilized by the technological behemoths.

The recently unveiled Harvard dataset, Institutional Books 1.0, is a monumental undertaking. It comprises nearly a million books, some dating back to the 15th century, written in a staggering 254 languages, encompassing over 394 million scanned pages.

While European languages like German, French, Italian, Spanish, and Latin still dominate, less than half of the volumes are in English, making it significantly more linguistically diverse than typical AI data sources.

The bulk of the collection hails from the 19th century, covering subjects from literature and philosophy to law and agriculture—all meticulously preserved and organized by generations of librarians. This treasure trove promises to be immensely beneficial for AI developers striving to enhance the accuracy and reliability of their systems.

“A lot of the data that’s been used in AI training is not from original sources,” observed Greg Leppert, executive director of the data initiative and chief technology officer at Harvard’s Berkman Klein Center for Internet & Society.

This new collection offers “down to the physical copy that was scanned by the institutions that actually gathered those materials,” providing a level of provenance rarely seen in previous AI training sets.

The concept of leveraging library archives for AI training isn’t entirely without precedent. Google, for instance, embarked on its ambitious Google Books project in 2006, scanning millions of volumes, including many copyrighted works, leading to protracted legal battles with authors that were only resolved in 2016.

Now, in a remarkable turn, Google is collaborating with Harvard to extract public domain volumes from its existing scans, paving the way for AI developers to access them.

The very group of authors who once sued Google, and more recently AI companies, now applauds this new initiative. “Many of these titles only exist on major library shelves, and the creation and use of this dataset will expand access to these volumes and the knowledge they contain,” said Mary Rasenberger, executive director of the Authors Guild, adding that it will “democratize the creation of new AI models.”

The collaboration also offers a lifeline to libraries facing the costly and meticulous task of digitalization. Jessica Chapel, director of digital and online services at the Boston Public Library, noted that when OpenAI first approached them, the library made it clear that any digitized information would be made publicly available.

“OpenAI had this interest in massive amounts of training data. We have interest in massive amounts of digital objects. So, this seems to be a case where interests are aligning,” Chapel explained.

Funding from tech giants can help finance projects librarians were eager to undertake anyway, such as scanning dozens of 19th and early 20th-century French-language newspapers from New England’s Canadian immigrant communities.

Yet, this embrace of historical data also introduces a nuanced set of challenges. A collection imbued with 19th-century thought, while invaluable for teaching AI to “plan and reason” by exposing it to “pedagogical materials on what it means to reason” and “scientific information on how to execute processes and how to execute analysis,” as Leppert points out, also contains outdated and potentially harmful content. Discredited scientific and medical theories, alongside racist and colonial narratives, are part of this historical record.

Kristi Mukk, coordinator at Harvard’s Library Innovation Lab, acknowledged these “complicated issues around harmful content and language.” The initiative is actively working to provide guidance to mitigate these risks, aiming to “help users make their own informed decisions and use AI responsibly.”

As these vast, nuanced datasets are shared on open-source platforms like Hugging Face, the true utility for the next generation of AI tools remains to be seen.

What is clear, however, is that AI is embarking on a new phase of its intellectual development. No longer content with the digital echo chamber of recent decades, it is now delving into the curated wisdom of the past, seeking a deeper, more authentic understanding of human experience.

This journey, from the ephemeral tweets to the weighty tomes, marks a significant evolution, one that brings both immense promise and the weighty responsibility of confronting the full, complex tapestry of human history.

Author

  • LNGFRM Team

    Frank DiBernardo handles LNGFRM's Foodie and Miscellaneous writing tasks. He's always getting ideas from users, so don't be afraid to send an email to the editor.

Daily Newsletter
Subscribe to our Newletter!
You May Also Like
© 2026 LNGFRM. All rights reserved.