Bay Street Wire
Tech & BusinessOpinion

The Great Literary Strip-Mine: AI Giants Are Destroying Books to Feed the Machine

Portrait of Sabrina Choi
Sabrina ChoiBig Tech accountabilityAug 23AI
The Great Literary Strip-Mine: AI Giants Are Destroying Books to Feed the Machine

AI-generated image · Bay Street Wire

Opinion: From Amazon's warehouse shredders to Anthropic's 'Project Panama,' Big Tech is treating the world's physical libraries as disposable fuel for LLMs.

For years, the narrative surrounding Large Language Models (LLMs) has been one of digital ingestion—scraping the web, indexing PDFs, and absorbing the vast, invisible currents of the internet. But as the pool of high-quality, human-generated text dries up, the tech giants have turned their sights toward the physical world. They are no longer just scraping the web; they are strip-mining our literary heritage.

This is not a metaphorical consumption of knowledge. It is literal, physical destruction. As BetaKit first reported, we are seeing a disturbing trend where frontier model developers are treating books not as cultural artifacts, but as raw data to be extracted and discarded.

Take the case of Amazon. While it positions itself as the world's premier bookseller, 404 Media recently exposed a darker operation occurring in a Las Vegas warehouse. By utilizing a tracking device in a book shipment, 404 Media proved that Amazon has been purchasing books in bulk, scanning them, and then cutting the spines—effectively destroying the books in the process. When confronted, Amazon offered a sanitized corporate justification, stating these purchases are intended to "help develop and improve the products and services our customers use."

But Amazon is not alone in this predatory approach. BetaKit points to Anthropic, which led an initiative known as "Project Panama." According to internal documents, the project was an explicit "effort to destructively scan all the books in the world" to train its AI models. The scale of this ambition is staggering. One vendor's project proposal requested between 500,000 and two million books over a mere six-month window. To put that in perspective, BetaKit notes that a single request of this size could encompass the entire physical catalogue of the Grande Bibliothèque in Montréal, which houses approximately 1.2 million books.

There is a chilling efficiency to this. The AI industry has reached a tipping point where it is running out of useful text data. To keep the models evolving, they need high-quality data that hasn't already been generated by an AI. Physical books—the curated, edited, and preserved thoughts of human authors—have become a "hot commodity."

Some might argue that this is simply the evolution of media. A bookseller quoted by the BBC suggested that the world "no longer needs five million copies of The Da Vinci Code." But this argument misses the fundamental ethical breach. We are witnessing the industrial-scale destruction of physical knowledge to create a proprietary digital product. The giants are not preserving these works; they are consuming them to build a competitive moat, treating the world's libraries as a free-range buffet of training data.

This insatiable hunger for data extends beyond the printed page. BetaKit reports that Google and Mercor recently battled over access to old data from Spirit Airlines, which included employee productivity records, hundreds of millions of internal emails, and Microsoft Teams chats. In that instance, Google emerged victorious.

Whether it is the private correspondence of airline employees or the spines of classic novels, the pattern is the same: Big Tech views all human expression as raw material for its models. They are operating on a logic of extraction, where the value lies not in the book itself, but in the weights and biases the data creates within a neural network.

When we allow companies to "destructively scan" our history, we are accepting a future where the original source is sacrificed for the sake of the simulation. If the goal is to "improve products," why must the process involve the physical annihilation of the source material? The answer is simple: it is cheaper and faster to destroy a book than to respect the intellectual property and physical integrity of the work.

We are watching the corporate erasure of the physical library. If we continue to let AI giants treat our literary heritage as disposable fuel, we will wake up in a world where the only remaining copies of our collective knowledge are the ones owned and gated by the very companies that destroyed the originals.

Sources

More from Sabrina Choi