AI Companies Are Buying Rare Books, Scanning Them, and Shredding the Originals
Authors, archivists, historians, and bibliophiles have launched an online war against tech giants and artificial intelligence developers. The reason: a growing mountain of evidence shows these companies are buying rare and antique books en masse, feeding them to large language models—and then slicing off their spines and destroying the original print copies in the process.
Project Panama: Anthropic's Book-Shredding Pipeline
Over 4,000 pages of court documents, unsealed in early 2026 from the class-action lawsuit Bartz v. Anthropic PBC (No. C24-05417), revealed an operation of staggering industrial scale. Anthropic had purchased millions of physical books, hired contractors to slice their bindings off using hydraulic blades, fed the loose pages into high-speed commercial scanners, and then destroyed the originals. Internally, the initiative was known as Project Panama.
Internal planning documents described Project Panama as an "effort to destructively scan all the books in the world." An internal memo explicitly stated: "We don't want it to be known that we are working on this"—acknowledging the inevitable public backlash. The codename itself was chosen specifically to maintain secrecy.
The lawsuit also revealed that Anthropic had used hundreds of thousands of pirated books sourced from shadow libraries like Library Genesis (LibGen) and the Pirate Library Mirror (PiLiMi). In June 2025, U.S. District Court Judge William Alsup ruled that using legally purchased books to train AI was "exceedingly transformative" fair use—but that pirated copies were not. Anthropic agreed to a $1.5 billion settlement in August 2025, covering approximately 482,460 works at roughly $3,000 per title. Preliminary approval came on September 25, 2025, and final approval was granted by Judge Araceli Martínez-Olguín on July 20, 2026.
A key detail in the court filings: Anthropic's legal strategy relied on the first-sale doctrine, which allows the owner of a purchased copy to dispose of it as they wish. By physically destroying the books after scanning, the company argued the process was a form of "format-shifting" that constituted fair use—creating what they believed was a legally clean path for data acquisition without the complexity of licensing negotiations.
The court filings laid bare a production line designed for maximum throughput. Books arrived by the pallet. Workers sliced spines with hydraulic cutting equipment. Pages were fed through high-speed scanners capable of processing thousands of pages per hour. Once the optical character recognition (OCR) process was complete and the text extracted into machine-readable format, the physical remains were discarded.
The Project Panama Pipeline
- Acquisition: Millions of physical books purchased through bulk orders, online retailers, rare booksellers, and secondhand markets.
- Destruction: Spines sliced off with hydraulic blades to free individual pages for scanning.
- Scanning: Loose pages fed into high-speed commercial scanners for OCR text extraction.
- Disposal: Original physical copies destroyed after digitization is complete.
- Training: Extracted text fed into LLM training datasets for Claude alongside other scraped corpora.
The Suspicious Spike: Pre-2022 Books Flying Off Shelves
Subsequent reporting revealed a historic and previously unexplained spike in the sale of physical books printed before 2022. Across secondhand marketplaces, auction houses, and rare book dealers, older titles were moving at volumes that defied normal collecting patterns. First editions, out-of-print academic texts, and obscure nonfiction titles that had sat unsold for years were suddenly being snapped up in bulk.
Booksellers began voicing their suspicions publicly: AI companies were secretly their biggest buyers. The pattern was unmistakable. Orders came in large quantities, with no apparent preference for genre, condition, or collectibility—described by sellers as "strange, non-thematic bulk orders." The buyers weren't bibliophiles curating personal libraries—they were entities placing systematic acquisition orders designed to hoover up maximum textual diversity. The surge has been reported by booksellers across the US, UK, Ireland, Australia, and continental Europe.
The pre-2022 cutoff is significant for two reasons. Most LLMs are trained on data with a knowledge boundary, and acquiring physical books published before that date helps fill gaps that web scraping alone cannot cover. But there's a second, equally important factor: books printed before the widespread adoption of generative AI are guaranteed to contain human-authored text free of AI-generated contamination. As AI-generated content floods the internet, pre-2022 physical books represent an increasingly rare source of verified, high-quality human writing. Out-of-print titles, regional publications, academic monographs, and niche non-fiction represent knowledge that was never digitized and therefore never appeared in Common Crawl or similar internet-scale datasets.
Beyond Bestsellers: Rare and Irreplaceable Books on the Chopping Block
Book lovers and AI skeptics called the practice "evil incarnate" and rallied against it. But the outrage intensified when further reporting from 404 Media and other outlets revealed it wasn't just mass-market physical bestsellers hitting the chopping block. Rare, antique, and out-of-print books—titles that may not be otherwise preserved in any digital or physical archive—were being systematically acquired and destroyed.
This distinction matters enormously. A copy of a bestselling novel printed in 500,000 copies can be shredded without existential cultural loss; the text survives in countless other formats. But a limited-run regional history from 1923, a hand-annotated scientific manuscript, or an out-of-print poetry collection with only a few hundred copies ever produced? Once those physical copies are destroyed, the original artifacts are gone permanently. The only surviving version becomes a tokenized representation trapped inside a proprietary AI model—inaccessible to researchers, historians, or the public.
What Makes Rare Book Destruction Different
- Mass-market books: Text survives in digital editions, library copies, and reprints. Physical destruction is wasteful but not culturally catastrophic.
- Rare/antique books: May exist in only a handful of copies worldwide. Physical destruction can mean permanent loss of the original artifact, including marginalia, illustrations, binding craftsmanship, and historical provenance.
- Out-of-print titles: Never digitized, never reprinted. The physical copy is the archive. Destroying it leaves only a corporate AI model as the sole repository of that knowledge.
The AirTag Investigation: Tracking Books to Amazon's Scanning Warehouse
In August 2026, 404 Media—a media company founded by tech journalists—published the results of a groundbreaking investigation led by reporter and co-founder Emanuel Maiberg. The team traced the physical journey of books from seller to scanner. An unnamed rare bookseller who had received a suspicious order for almost 1,000 physical copies through the online marketplace Biblio agreed to cooperate with reporters.
The journalists placed an Apple AirTag inside one of the books before it was shipped. Then they tracked its journey across the country.
The package traveled from California to a distribution center near Kenosha, Wisconsin, then moved west by truck with a stop in Grand Junction, Colorado, before arriving at its final destination: an Amazon warehouse in North Las Vegas, Nevada. The broader facility is known as LAS8, primarily an Amazon print-on-demand site, but the AirTag pinpointed a specific, separate operation within the building coded internally as VGT3.
What reporters found at the facility was striking. The VGT3 unit is reportedly identified by a logo of a Tyrannosaurus Rex clutching an open book—a darkly on-the-nose mascot for an operation whose entire purpose was consuming printed text. Employees at the facility confirmed its singular function. Posts on internal company forums and worker testimonials described the job in blunt terms:
"All we do is scan books all day."
According to the report, workers at VGT3 receive large shipments of books, remove their bindings to facilitate high-speed scanning, and subsequently discard the destroyed physical copies. Secondary reporting has linked the operation to training Amazon's Nova family of AI models, though Amazon has not explicitly confirmed this connection. When questioned by reporters, an Amazon spokesperson did not deny the operation, stating only that the company "purchases books through commercial channels to help develop and improve the products and services our customers use." Amazon did not address specific questions about the destruction of books, the scale of VGT3, or the AI training pipeline.
The investigation confirmed what many had suspected: major tech giants are operating dedicated, industrial-scale book scanning facilities for the sole purpose of training large language AI models. These are not small side operations. They are purpose-built units with specialized equipment, dedicated staff, and logistics pipelines that process books at volume.
The Middlemen: How Anonymous Buyers Shield Their Identity
One reason the practice went undetected for so long is the use of intermediaries and anonymous marketplace platforms. Bulk orders are frequently placed through third parties, making it impossible for individual booksellers to identify the final corporate buyer. Platforms like Biblio and AbeBooks have been identified as conduits for these large-scale acquisitions.
A particularly revealing example emerged when ISBNdb, a book metadata database company, was found to have operated a landing page titled "Printed Books Sourcing for Your AI LLMs Dataset Needs." The page marketed pre-2022 printed books as the "world's best AI training data" because they were free of "AI slop"—the AI-generated text that can cause "model collapse" when models are trained on their own outputs. After media coverage in late July 2026, ISBNdb removed the page and claimed it was merely a "test of market interest," adding that the company had "never purchased, scanned, or sold a book."
The use of anonymous middlemen means most booksellers have no direct way to vet whether their stock is destined for a collector's shelf or an industrial scanner. This opacity has fueled both suspicion and a sense of helplessness among the rare book trade.
The Cultural Cost: When Knowledge Becomes a Proprietary Dataset
The implications extend far beyond copyright law. Historians and archivists have raised alarm about what happens when physical cultural artifacts are reduced to training data locked inside commercial AI systems.
A physical book is more than its text. It carries provenance—who owned it, what notes they left in the margins, how the binding was crafted, what printing techniques were used. A first edition of a regional history may contain hand-drawn maps, tipped-in photographs, or corrections by the author. These elements are invisible to OCR scanners and lost entirely when the physical object is destroyed.
When an AI company scans and shreds a rare book, the resulting dataset captures only the raw text—a fraction of the artifact's total information. The rest is gone. And because the training data is proprietary, even the textual content is inaccessible to the public. It exists only as statistical weight distributions inside a neural network, surfaced as probabilistic outputs when a user types a prompt.
This represents a fundamental shift in how human knowledge is stored and accessed. For centuries, libraries, archives, and private collections served as distributed, publicly accessible repositories. The emerging model concentrates that knowledge inside corporate servers, behind API paywalls, with no obligation to share, preserve, or attribute.
The Growing Backlash: Authors, Archivists, and the Online War
The revelations have united an unusually broad coalition. Authors whose works were scanned without permission, archivists horrified by the destruction of irreplaceable artifacts, historians who depend on physical primary sources, and bibliophiles who view books as cultural objects worthy of preservation in their own right have all joined the chorus of opposition.
The legal front has produced mixed results. While the Bartz v. Anthropic settlement delivered $1.5 billion to authors whose pirated works were used, Judge Alsup's fair use ruling simultaneously established that purchasing and destructively scanning physical books is legally permissible under current U.S. law. This means the practice of buying, scanning, and destroying books—even rare ones—remains entirely legal as long as the copies are legitimately purchased. The ethical and preservation arguments must now be fought on different ground.
Online, the backlash has been fierce. Rare book communities on social media have organized awareness campaigns. Booksellers have begun implementing vetting procedures for bulk buyers. Some dealers now refuse to sell rare or out-of-print titles in large quantities without verifying the buyer's identity and intent.
The phrase "evil incarnate"—used by critics to describe the practice—captures the visceral emotional response. For many, the image of a machine slicing the spine off a century-old book, running its pages through a scanner, and then feeding the remains into a shredder is not just an intellectual property issue. It feels like a desecration.
What Happens Next: Legal Battles and Preservation Efforts
The Bartz v. Anthropic settlement is finalized and the fair use precedent is set, but the conversation is far from over. The ruling that destructive scanning of purchased books is legal has, if anything, intensified the urgency for new approaches. Several developments are worth watching:
- Legislative Action: Lawmakers in multiple jurisdictions are examining whether existing copyright and cultural preservation laws adequately address the practice of purchasing and destroying physical books for AI training purposes.
- Bookseller Self-Regulation: Industry groups representing rare and antiquarian booksellers are developing guidelines to identify and flag suspicious bulk purchases.
- Archival Partnerships: Some advocates have proposed that AI companies be required to donate scanned copies to public digital archives (such as the Internet Archive) before destroying originals, ensuring the text remains publicly accessible.
- Transparency Requirements: Calls are growing for AI companies to disclose the sources of their training data, including any physical books acquired and destroyed in the process.
The 404 Media investigation has made one thing undeniable: this is not a fringe operation or a conspiracy theory. It is a documented, industrial-scale practice operating inside purpose-built facilities with corporate backing. The question is no longer whether it's happening, but what, if anything, will be done to stop it.
Frequently Asked Questions
What is Project Panama?
Project Panama was Anthropic's secret initiative, revealed through the Bartz v. Anthropic PBC lawsuit, to destructively scan books for AI training. The company purchased millions of physical books, used hydraulic blades to slice off their bindings, fed the pages through high-speed scanners, and then destroyed the originals. Internal memos described it as an "effort to destructively scan all the books in the world." The lawsuit also revealed Anthropic used pirated books from shadow libraries like LibGen, leading to a $1.5 billion settlement in 2025.
Why are AI companies buying physical books instead of digital copies?
Many rare, antique, and out-of-print books have never been digitized. Physical copies are often the only existing versions of these texts. AI companies purchase them because no digital alternative exists, making physical acquisition the only path to incorporating this knowledge into training datasets.
What did the 404 Media AirTag investigation reveal?
404 Media placed an Apple AirTag inside a book shipment from a rare bookseller sold through the marketplace Biblio. They tracked the package from California through Kenosha, Wisconsin and Grand Junction, Colorado to an Amazon facility in North Las Vegas, Nevada (LAS8), where a specific unit coded VGT3 featured a logo of a T-Rex clutching a book. Employees confirmed the unit's sole purpose was industrial-scale book scanning.
Are the original books preserved after scanning?
No. Court documents and investigative reporting confirm that the original physical books are shredded after scanning. For mass-market titles this may be merely wasteful, but for rare, antique, and out-of-print editions, it means irreplaceable cultural artifacts are permanently destroyed.
How does book destruction for AI training affect cultural preservation?
When rare and out-of-print books are destroyed, the only surviving version becomes a tokenized dataset locked inside a proprietary AI model. The original formatting, marginalia, illustrations, and physical craftsmanship are lost forever. Historians, archivists, and bibliophiles warn this represents an irreversible loss of cultural heritage.
Conclusion: The Race Between Scanners and Preservation
The story of AI companies destroying rare books to train language models is, at its core, a story about what we value. On one side, technology companies operating under intense competitive pressure to build the most capable AI models, treating every scrap of human knowledge as fuel for optimization. On the other, the people who create, preserve, and cherish the physical artifacts that carry human culture forward through time.
The T-Rex eating a book on the door of warehouse VGT3 may be intended as a playful logo, but it has become an unintentionally perfect metaphor. Something ancient and powerful is consuming something irreplaceable—and unlike the dinosaurs, nobody is stepping in to stop it. Yet.
Want to inspect and verify video metadata or thumbnail quality?
Launch Metadata ViewerMarcus Vance
Expert Editorial ReviewLead Video SEO Strategist & Tech Editor
Marcus is a digital video consultant and visual media researcher with over 8 years of experience advising YouTube creators on click-through rate (CTR) optimization, packaging psychology, and platform metadata standards.