The Great Data Heist: AI Giants Are Destroying Books to Feed the Machine

AI-generated image · Bay Street Wire
From Amazon's warehouse shredders to Anthropic's 'Project Panama,' Big Tech is strip-mining physical libraries to solve its looming data drought.
The 'information age' is rapidly evolving into a corporate heist of human creativity. As the world's frontier model developers exhaust the supply of high-quality, human-generated text available online, they have turned their sights toward the physical world—and they are leaving a trail of destruction in their wake.
As BetaKit first reported, the insatiable hunger for training data has led AI companies to target physical books that were never digitized. This isn't just a matter of procurement; it is a matter of systemic destruction. BetaKit reports that 404 Media recently exposed Amazon for buying books in bulk and transporting them to a warehouse in Las Vegas. There, the company has been scanning the volumes and cutting the spines, effectively destroying the books in the process. When questioned, Amazon stated that these purchases are intended to "help develop and improve the products and services our customers use."
Amazon is far from alone in this predatory approach. BetaKit notes that Anthropic led a similar initiative known as "Project Panama." According to internal documents, this was an "effort to destructively scan all the books in the world" specifically to train AI models. The scale of these ambitions is staggering: one vendor's project proposal requested between 500,000 and two million books over a six-month window. To put that in perspective, BetaKit points out that a single request of that size could encompass the entire physical catalogue of the Grande Bibliothèque in Montréal, which houses approximately 1.2 million books.
This desperation for "clean" data—meaning content not generated by AI—has extended beyond the library stacks. BetaKit reports a recent battle between Google and Mercor over access to legacy data from Spirit Airlines. The dispute centered on hundreds of millions of internal records, including employee productivity logs, Microsoft Teams chats, and internal emails. Google ultimately won the rights to this data.
While some, such as a bookseller quoted by the BBC, argue that the world no longer needs millions of copies of popular titles like *The Da Vinci Code*, the broader implication is clear. The giants of the industry are no longer content to crawl the open web; they are now treating the world's intellectual heritage as raw material to be harvested, processed, and discarded. By treating books as disposable fuel for their LLMs, these companies are not just innovating—they are erasing the physical artifacts of human thought to build a corporate mirror of it.

