When AI Destroys Books to Learn From Them
The idea of destroying rare books to feed artificial intelligence sounds like something from a dystopian novel. Recent reports suggest that some AI companies are doing exactly that—scanning, digitizing, and then physically destroying rare or out-of-print volumes to build training datasets. The practice raises uncomfortable questions about the cost of progress, the value of cultural heritage, and whether the rush to build ever-larger language models justifies irreversible loss.
At the center of this controversy is a model called Kimi-K3, released on HuggingFace earlier this year. According to its technical report, the model was trained on a massive corpus of text that includes digitized versions of rare books, some of which are no longer in print and exist in only a handful of physical copies worldwide. After scanning, the original volumes were reportedly shredded to prevent redistribution or resale, a step taken to comply with certain licensing interpretations or to avoid legal complications around copyrighted material.
This approach isn’t entirely new. For years, companies have scanned books for projects like Google Books, though those efforts typically preserved the originals. What’s different now is the scale and the finality. AI training demands vast quantities of text, and rare books—especially those with unique dialects, archaic language, or niche subject matter—can offer linguistic diversity that improves model performance. But when the only copy of a 19th-century regional diary or a limited-run poetry chapbook is destroyed after scanning, the loss isn’t just financial. It’s historical.
Critics argue that this treats cultural artifacts as disposable inputs rather than irreplaceable objects. Librarians and archivists have voiced concern that the drive for training data is undermining decades of preservation work. One rare book curator, speaking on condition of anonymity, said, “We’re not just losing paper and ink. We’re losing provenance, marginalia, water stains, the physical evidence of how a book was used and loved. That context matters.”
Supporters of the practice point out that many of these books were already deteriorating or lacked clear ownership records. In some cases, scanning and shredding was done with the cooperation of institutions that lacked the resources to preserve the materials themselves. They also note that digital preservation, even if it requires destroying the original, can make content accessible to researchers worldwide in ways that a locked archive never could.
Still, the ethical line remains blurry. Unlike public domain works, many of the books in question may still be under copyright, even if commercially unavailable. The technical report for Kimi-K3 acknowledges that the training data includes copyrighted material but claims reliance on fair use arguments—a position that remains untested in court for AI training at this scale.
The broader implications extend beyond books. If AI companies begin treating other forms of media—vinyl records, film reels, handwritten manuscripts—as raw material to be consumed and discarded, the cultural cost could grow significantly. Some compare it to the early days of the internet, when content was scraped and repurposed with little regard for origin or intent. Others see it as a symptom of an AI bubble prioritizing speed and scale over responsibility.
There are alternatives. Projects like the Internet Archive and various university-led initiatives have shown that mass digitization can happen without destruction. Collaborative models, where AI developers pay for access to digitized collections through libraries or licensing agreements, could offer a path forward. Some startups are even exploring synthetic data generation to reduce reliance on real-world corpora altogether.
For now, the sight of a rare book being fed into a shredder after its final scan serves as a stark metaphor. It reminds us that every advance in AI is built on choices—about what we value, what we sacrifice, and who gets to decide. As models like Kimi-K3 push the boundaries of what machines can understand, we might do well to ask not just what they can learn, but what we’re willing to lose in the process.
The future of AI doesn’t have to be built on the ashes of the past. But if we keep treating cultural heritage as fuel, we may find ourselves with remarkably intelligent machines—and far less to remind them why intelligence matters.
