Hook
Anthropic spent millions buying physical books—then systematically destroyed them. The paper was scanned, the spines cracked, the originals shredded. This wasn't censorship. It was a data acquisition strategy. And it’s the most legally creative arbitrage I’ve seen since the 2017 ICO days.
Context
The AI industry is starving for clean, human-generated text. Web-scraped data is polluted with AI-generated content. Copyright lawsuits are piling up. Traditional licensing is slow and expensive. So a new solution emerged: buy physical books, scan them destructively, and shred the originals. The legal basis? A 2025 US court ruling that converting a lawfully purchased physical book into a non-distributable digital copy—as long as the original is destroyed—qualifies as fair use. It’s called the “one-to-one replacement” doctrine.
Core
The execution is brutal and efficient. According to sources, Anthropic hired the former head of Google’s book scanning project. They partnered with ISBNdb, a company that specializes in purchasing and destructively scanning books. ISBNdb’s marketing material explicitly states they buy books by ISBN, filter by publication year, and can “verify the destruction” through legally binding NDAs. The pitch: pre-2022 books are less likely to contain AI-generated text or data poisoning artifacts, making them “pure” training material.
But here’s the catch: once scanned, the physical book is thrown away. Not archived. Not donated. Shredded. The logic is simple—if you keep the digital copy, you can’t keep the physical one without violating copyright. The court reasoned that replacing a physical copy with a digital one, without increasing the total number of copies, is transformative fair use.
This isn’t a small experiment. Anthropic spent “millions of dollars” on “millions of books.” The volume is staggering. ISBNdb now offers this as a standard B2B service for AI developers. They boast “exclusive” access to rare, out-of-print titles—because once they’re scanned and destroyed, no one else can legally access them. It’s a data moat built on physical scarcity.
The financial engineering here is clean: pay retail price for a book, scan it, destroy it, and own the exclusive digital rights to its text forever. The per-book cost is trivial compared to the value of training data. But the hidden cost is legal tail risk. The 2025 ruling only covered non-distributive library copies. Once those scans are fed into a large language model and served to millions of users, is that “distribution”? The lawsuit against Anthropic for allegedly pirating Library Genesis copies is still pending. That’s the crack in the foundation.
Contrarian
Let’s call this what it is: data arbitrage, not innovation. The court’s “one-to-one replacement” logic assumes digital copies behave exactly like physical ones. They don’t. A digital file can be replicated infinitely with zero marginal cost. The moment you create that scan, you have the technical ability to copy it a million times. The legal constraint relies entirely on enforcement. And enforcement is weak.
What happens when a disgruntled employee copies the training corpus? What happens when a cloud backup leaks? The physical book is gone. The digital copy is fragile but infinitely reproducible. This isn’t preservation—it’s a legally sanctioned destruction of cultural artifacts. The article rightly notes that no one has disclosed which specific rare books were destroyed. That silence is deafening.
Moreover, this model creates an unsustainable arms race. Physical books are a finite resource. The total stock of pre-2022, high-quality, human-authored texts is limited. As more AI companies adopt this strategy—and they will—prices will spike, and the low-hanging fruit will vanish. The real moat isn’t the data; it’s the legal fiction that lets you destroy it first. Speed is the only currency that doesn’t depreciate, but here speed destroys the resource itself.
Takeaway
The bet here is that the 2025 ruling holds. If it does, we’ll see a gold rush on physical libraries—warehouses full of books bought, scanned, and incinerated. If it doesn’t—if an appeals court or Congress closes the loop—the entire model collapses, and those millions of dollars become sunk cost. We don't predict the market—we predict where people will look for answers. Right now, they’re looking at court dockets, not bookstores.