Key takeaways
- ISBNdb deleted its AI data service after reports revealed AI firms were bulk-purchasing rare books to destroy for model training material.
- Anthropic settled for $1.5 billion over Project Panama, an internal program that scanned and destroyed books in bulk; the practice was legal but generated massive negative publicity.
- AI companies face a genuine shortage of high-quality training data, pushing them toward rare books—a crisis unlikely to resolve until transparent licensing agreements become standard practice.
ISBNdb, an online book database, has shuttered an artificial intelligence service following widespread outrage over revelations that AI firms were bulk-purchasing rare books and destroying them to fuel model training. The company’s deleted landing page had positioned itself as a middleman, connecting AI developers directly to book suppliers while promising access to “curated, peer-reviewed, domain-specific human knowledge” that surpassed standard web-scraped material. The arrangement faced immediate backlash once the practice became public.
ISBNdb responded with damage control. “We’ve seen the recent coverage about a marketing landing page on our site, and we understand the concern it raised,” the company stated. “We don’t train AI models, and we never have. The page was a test of market interest; no such service was ever brought to life. We’ve taken the page down.”
The now-deleted landing page had described books using language designed to appeal to AI developers struggling with data quality. “The world’s best AI training data is sitting on a shelf,” it read. “Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative.” These claims tapped into a genuine crisis facing the AI industry: the increasing difficulty of sourcing high-quality training material that could produce reliable, sophisticated models.
Investigators Uncovered Unusual Purchasing Patterns
The practice came to light when 404 Media began investigating suspicious bulk purchasing behavior among rare book suppliers. Schools and libraries, already stretched thin by budget constraints, suddenly found themselves contacted with offers to purchase large collections at premium prices. These institutions, lacking resources to match the offers, often sold materials they had acquired over decades.
Suppliers Faced Conflicting Incentives
One anonymous book seller described the moral conflict to 404 Media reporter Samantha Cole. “It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell,” the supplier said. “On the other hand, I don’t like the end-use, and I don’t like that uncommon books are being pulped.” The comment captured the central tension of the scheme: legitimate financial benefit for struggling institutions, paired with knowledge that rare materials faced destruction rather than preservation.
Why Educational Institutions Became Targets
Schools and libraries operate under severe budget constraints, making them unable to resist offers that significantly exceeded what they might expect to earn through traditional resale. Rare book markets move slowly; a medieval manuscript or out-of-print academic text might languish on shelves for years. Bulk offers provided immediate liquidity during a period when many institutions faced facility closures and program cuts.
Project Panama Revealed the Systematic Scope
The full scale of the operation became apparent in January when unsealed court documents exposed Project Panama, an internal Anthropic program designed to scan and destroy books in bulk for language model training. The revelation proved damaging enough that Anthropic paid authors $1.5 billion in settlement.
The most troubling aspect: the court filings indicated the destruction was technically legal. Anthropic had proceeded with full knowledge that bulk book destruction would generate significant negative publicity, but calculated that settlement represented a more economical path than sourcing alternative training material or accepting lower model performance.
Legality Does Not Preclude Expensive Consequences
The distinction mattered. No laws explicitly prohibited Anthropic from purchasing books and destroying them as part of the training process. The legal vulnerability lay not in the activity itself, but in its reputational impact. The company determined that $1.5 billion in settlements cost less than the long-term damage from continued public association with the practice.

The Underlying Data Crisis
These practices emerged from a genuine problem facing artificial intelligence developers: the shortage of high-quality training material. The explosion of AI model development has exhausted readily available, legitimate sources of premium training data.
Web Content Falls Short for Sophisticated Models
Data pulled from internet crawls presents multiple problems. Generic web content contains redundancy, inconsistency, factual errors, and bias. Early generations of large language models trained on unfiltered internet material produced outputs riddled with hallucinations and contradictions.
More troubling, researchers began documenting “model collapse”—a phenomenon where AI-generated content contaminates training datasets, causing subsequent models trained on that data to degrade further. When AI trains on AI-generated material, output quality systematically declines across generations. This created pressure to source genuinely original, human-created material.
Why Books Represent a Premium Asset
Published books offered an alternative. Professional publishers subject manuscripts to editorial review, fact-checking, copyediting, and revision cycles. A book released by a major house theoretically represents material screened for accuracy and quality. Rare and specialized texts—first editions, technical manuals from defunct fields, academic treatises on obscure subjects—contained knowledge unavailable through any other source.
This gap between data quality needs and legitimate supply created the economic conditions where schemes like Anthropic’s Project Panama seemed reasonable from a business perspective.
Corporate Partnerships Signal a Shift
A24, the entertainment company, announced a summer partnership with Google supporting training operations for DeepMind, Google’s AI research division. The collaboration positioned itself as a formal media licensing agreement rather than a covert acquisition program, attempting to legitimize data sourcing through explicit contracts and compensation structures.
The move signaled that at least some major technology firms recognized the reputational cost of opaque sourcing arrangements and were willing to pursue transparent partnerships as an alternative.
Cost-Cutting Drove Physical Destruction
The unsealed Project Panama documents did not suggest rare book destruction served some strategic purpose of knowledge suppression or collection control. More simply, destruction reflected economic decisions made during the scaling phase of data acquisition.
Book digitization need not destroy originals. Proper scanning equipment and careful handling preserve physical material while capturing its contents for training purposes. However, preservation requires investment—climate-controlled storage, trained handling, archival materials. Destruction represented the cheaper alternative: rapid scanning using faster methods that damaged spines and pages, then discarding the compromised material.
The choice between preservation and destruction hinged entirely on whether companies prioritized long-term cultural value against immediate cost savings. The decision to destroy reflected priorities: the material’s value lay only in its training utility, not in its existence as a historical artifact.
Implications for Future AI Development
ISBNdb’s service shutdown and the public fallout from Project Panama signal growing scrutiny on AI data sourcing practices. Authors, publishers, and institutions are now alert to suspicious purchasing patterns. The initial outrage may deter future similar operations.
Yet the underlying tension remains unresolved. AI firms still require high-quality training material, and legitimate sources remain constrained. Whether the industry moves toward transparent licensing structures like A24’s partnership, develops open-source datasets, or discovers entirely new sourcing models will determine whether book-based training material plays a role in future AI development.
Frequently Asked Questions
What was ISBNdb's AI service?
ISBNdb offered a now-deleted landing page marketing itself as a middleman connecting AI firms with book suppliers to access "curated, peer-reviewed" training material positioned as superior to web-scraped data.
Why did Anthropic pay $1.5 billion?
Court documents exposed Project Panama, Anthropic's program to scan and destroy books in bulk for training. Though technically legal, the practice generated such negative publicity the company settled rather than face continued reputational damage.
Why were the books destroyed instead of preserved?
Cost-cutting drove the practice. Cheaper, faster scanning methods damaged books beyond preservation; companies chose destruction over investing in archival storage, handling, and proper preservation processes.