AI's Quiet War on Cultural Heritage: How Rare Books Are Being Destroyed to Train Language Models
There’s a quiet crisis unfolding in the back rooms of some AI labs and data centers. Rare books—centuries-old manuscripts, out-of-print scientific treatises, handwritten letters from forgotten scholars—are being fed into industrial shredders. Not as part of some dystopian art project, but as a routine step in training the next generation of large language models.
The irony is hard to miss: to build systems that can understand and generate human language with uncanny fluency, companies are destroying the very artifacts that embody centuries of human thought.
This isn’t happening in the open. No press releases announce the pulping of 18th-century botanical guides or first editions of early computing texts. But whistleblowers, librarians, and digital archivists have begun to speak up. They describe shipments of donated or acquired collections arriving at AI training facilities, only to be disassembled page by page, scanned for text, and then discarded.
The reasoning, when offered, is pragmatic: rare books often contain unique vocabulary, archaic syntax, and domain-specific knowledge not well represented in modern corpora. Including them, the argument goes, helps models handle edge cases, historical texts, or specialized jargon.
But the cost is steep. Many of these volumes are irreplaceable. A single copy of a 17th-century alchemy manuscript might be the only surviving record of a forgotten experimental method. Once shredded, that knowledge vanishes—not just from library shelves, but from the cultural record. Unlike digital files, which can be backed up and shared infinitely, a physical book destroyed in a shredder is gone forever. And while companies may scan the pages before destruction, the act of scanning often damages fragile bindings or ink, and the resulting OCR text is frequently riddled with errors, especially in older fonts or marginalia.
Some defenders of the practice point to the greater good: if feeding a few dozen rare books into a shredder helps create an AI that can assist researchers, translate ancient scripts, or democratize access to knowledge, isn’t it worth it? But critics counter that this logic treats cultural heritage as raw material to be consumed, not preserved. They note that many of these books were donated under the assumption they’d be housed in archives or made available to scholars—not fed into a machine learning pipeline and then thrown away.
There are alternatives, of course. Non-destructive scanning technologies exist that can capture text and images without harming the original object. Projects like the Internet Archive and various university-led digitization initiatives have shown it’s possible to preserve both the intellectual content and the physical artifact. But these methods are slower, more expensive, and require careful handling—qualities that don’t always align with the fast-moving, scale-obsessed ethos of AI development.
The issue raises broader questions about how we value knowledge in the age of AI. Are we willing to sacrifice the past to build a smarter future? And who gets to decide what counts as expendable? For now, the shredders keep running, their noise drowned out by the hum of servers training the next model. But somewhere, in a quiet corner of a library, a curator is wondering whether the next book to go might be the one that holds the key to understanding not just language, but ourselves.
