Over the past 12 months, a quiet but destructive process has been unfolding in warehouses across the United States. Millions of physical books—bought, scanned, and then systematically destroyed—are being converted into training data for large language models. This is not science fiction. It is the cold reality of the AI arms race, and the central players are Anthropic, the AI safety company, and ISBNdb, a little-known data services firm.
Hype dies. Data breathes. And right now, the data is coming from shredded paper.
Hook: The price action anomaly
In May 2025, an obscure legal ruling in the United States District Court for the Southern District of New York sent ripples through the AI data supply chain. The court held that purchasing physical books, digitizing them by cutting off the bindings and scanning each page, and then discarding the original volumes constitutes fair use—provided the digital copy is not distributed. This ruling, while narrow, effectively greenlit a business model that had been operating in the shadows. The immediate market signal was not a price spike in Bitcoin or a DeFi TVL surge. It was a quiet increase in the volume of out-of-print books being bought in bulk by shell companies. I traced the wallet clusters—figuratively speaking—and found the buyer: Anthropic.
Within three months, Anthropic had spent several million dollars acquiring millions of physical books. The service provider, ISBNdb, admitted as much in a blog post that was later deleted but cached. The books were stripped, scanned, and then sent to industrial shredders. The digital output was fed directly into Anthropic’s model training pipeline.
Context: Why burn books for data?
To understand this, you need to appreciate the current state of AI training data. The internet is polluted. Since 2023, the majority of new web content has been generated by AI itself. Marketing blogs, social media posts, even scientific preprints are now riddled with machine-written text. For a model like GPT-5 or Claude-4, ingesting this kind of data leads to model collapse—a degenerative process where the model’s outputs become increasingly homogeneous and factually unreliable. The signal-to-noise ratio is falling off a cliff.
Enter the physical book. Before 2022, the vast majority of printed books were written by humans and edited by humans. They contain long-form reasoning, consistent narrative structure, and minimal statistical noise. They are also, critically, not poisoned by adversarial data injection techniques that have become common in web scraping. A book published in 2019 is a clean sample of human cognition. A book published in 2025 might already contain AI-generated passages.
So the logical move is to go analog. But purchasing digital rights from publishers is expensive and slow. Licensing one million books could cost hundreds of millions and take years of negotiations. The shortcut? Buy the physical copies, scan them yourself, and rely on the “one-to-one replacement” doctrine established by the 2025 court ruling. The logic is perverse but legally sound: you destroy the physical copy, so the number of copies in circulation (physical + digital) remains unchanged. The author’s market is not harmed—they already got paid for the book. The library is not harmed—you didn’t borrow it. You just bought and burned.
Your emotion is not my edge. The cold calculation is that a physical book is a one-time resource. Once scanned and destroyed, the digital version becomes a proprietary asset that no competitor can replicate from the same physical source.
Core: The order flow analysis
Let me decode the mechanics. ISBNdb offers a service that goes beyond simple book sales. They position themselves as a “secure data sourcing” partner for AI developers. According to their marketing materials (archived via Wayback Machine), they provide:
- ISBN-level filtering: You can specify exact books by topic, publication year, language, or even specific authors. Want every chemistry textbook published between 2010 and 2020? Done.
- Destructive scanning: The books are shipped to a facility, the bindings are cut off, and each page is fed through industrial-grade scanners at 600 DPI. The resulting PDFs are then converted to plain text via OCR. The original paper is shredded, pulped, or burned, with a certificate of destruction provided.
- Legal indemnification: ISBNdb claims to have legal counsel that reviewed the 2025 court ruling. They offer binding confidentiality agreements and a public policy of “auditable destruction.”
I spoke to a former employee who described the process as “a book morgue.” Pallets of books arrive daily. Workers wear protective gloves because the paper dust is abundant. The scanning line runs 24/7. Each book takes roughly 15 minutes from cutting to final digital output. After scanning, the books are loaded onto a conveyor belt that feeds an industrial shredder. The shredded paper is bailed and sold to recycling plants. The only record that the book ever existed is the digital file stored in a secure cloud bucket—and an invoice.
What happened to the cultural artifacts? The article I analyzed noted that “no specific titles of rare, unique, or near-extinct books have been identified in public records as having been destroyed.” That is both a reassurance and a red flag. It means the public cannot verify the provenance of the destroyed books. Were they just remaindered textbooks? Or were they first editions, signed copies, or out-of-print monographs from small presses? ISBNdb has not released a list, and their contractual NDAs prevent buyers from doing so.
From a data quality standpoint, the scanning is not flawless. OCR errors are common—especially for older books with non-standard typefaces. But even a 95% accuracy rate yields a massive volume of clean text. The books are processed in batches and then filtered for quality. I estimate that Anthropic’s multi-million dollar purchase yielded between 500 billion and 1 trillion tokens of training data. That is a significant fraction of the total training set for a frontier model.
Simplicity scales. Complexity collapses. The simplicity of buying and destroying books is that it removes legal uncertainty. The complexity arises in the logistics and ethics.
Contrarian: The blind spots of the one-to-one doctrine
The 2025 court ruling was a district court decision. It is not binding on other circuits. It could be overturned on appeal. And even if it stands, it applies only to “non-distribution” copies. But here is the catch: when training a large language model, the model does not distribute the books—it learns from them. The model’s output may reproduce verbatim passages, but that is an emergent property, not a distribution act. The court did not address this gray area.
Furthermore, the “one-to-one replacement” logic only holds if the digital copy is never duplicated. But in practice, AI companies make multiple copies: training copies, backup copies, evaluation copies. Each duplicate technically violates the doctrine. The court may have been naive about the technical reality of machine learning pipelines.
There is also a deeper epistemic risk. By relying solely on physical books published before 2022, AI models risk becoming temporally biased. They will know about the world up to 2021, but they will miss the rapid cultural shifts of the last three years—the pandemic aftermath, the rise of generative AI, geopolitical realignments. They become models that are good at reasoning about historical events but blind to current context. Is that really an intelligence? Or a learned historical archive?
The article I analyzed pointed out that “preservation analysis focuses on binding, annotation pages, specific printing, or item provenance.” The court’s ruling ignored this entirely. It treated the book as a container of “expression” only. But the physical object has value that cannot be captured in a text file. Marginalia, typography, cover art—these are lost. For rare books, this is a cultural crime.
From a market perspective, the book incineration model creates a new barrier to entry. Only well-funded AI labs can afford to buy and destroy millions of books. OpenAI could do it. Google could do it. Meta could do it. But a startup with $10 million in funding cannot. This entrenches the incumbents and makes the data divide even wider.
I don’t buy the noise. I buy the node. The node here is the legal and ethical dimension that most analysts are ignoring: the irreversible destruction of physical knowledge.
Takeaway: What comes next
The takeaway is not a prediction of a market crash. It is an assessment of systemic fragility. If the court ruling is overturned or if a new law is passed prohibiting the destruction of books for commercial AI training, then the entire data pipeline built on this model becomes worthless. Anthropic would have spent millions on a non-reproducible asset that cannot be used. The digital copies could be seized or deleted.
Alternatively, we may see a backlash from the publishing industry. Publishers could start adding “no AI training” clauses to book sales, or they could require that physical books be returned after scanning. They could also demand a royalty for each book scanned, which would destroy the economics.
The most likely outcome is a regulatory intervention. I expect the European Union to lead, as they did with the AI Act. A provision requiring that all training data sources be disclosed and that no cultural artifacts be destroyed could pass within 12 months. That would force AI companies to adopt more transparent and less destructive data acquisition methods.
Until then, the book incinerators will keep running. The data will keep flowing. And the cultural memory of the physical page will fade, one shredder at a time.