The Great Literary Strip-Mine: AI Giants Destroy Books to Feed the Machine

AI-generated image · Bay Street Wire
From Amazon's Las Vegas warehouse to Anthropic's 'Project Panama,' frontier model developers are treating physical libraries as raw material for training data.
The race for artificial intelligence dominance has evolved into a literal war of attrition against the world's printed word. As reported by BetaKit, frontier model developers are exhausting the available high-quality text data on the internet and turning their sights toward physical books—not as treasures to be preserved, but as raw materials to be processed and discarded.
Reporting from BetaKit reveals a disturbing trend of 'destructive scanning' where AI companies are treating the physical remnants of human literary heritage as mere speed bumps. In a recent investigation by 404 Media, a tracking device was used to prove that Amazon has been purchasing books in bulk and transporting them to a warehouse in Las Vegas. There, the company has been scanning the texts and cutting the spines, effectively destroying the books in the process. When questioned, Amazon stated that these purchases are intended to "help develop and improve the products and services our customers use."
Amazon is not alone in this predatory approach to data acquisition. BetaKit reports that Anthropic led a similar initiative known as "Project Panama." Internal documents from the company describe this as an "effort to destructively scan all the books in the world" to provide training data for its AI models.
The scale of this appetite is staggering. BetaKit notes that one vendor's project proposal requested between 500,000 and two million books over a six-month window. To put this in perspective, a single request of this magnitude could encompass the entire physical catalog of the Grande Bibliothèque in Montréal, which is estimated to house 1.2 million books.
This desperation for non-AI-generated, high-quality data is extending beyond the library stacks. BetaKit reports that Google and Mercor recently battled over access to legacy data from Spirit Airlines, which included employee productivity records, hundreds of millions of internal emails, and Microsoft Teams chats. Google ultimately won the dispute.
While some, such as a bookseller quoted by the BBC, argue that the world does not need millions of copies of popular titles like *The Da Vinci Code*, the systemic destruction of physical texts represents a chilling shift. The giants of Big Tech are no longer content to index the world's information; they are now strip-mining it, destroying the physical evidence of human authorship to fuel the next iteration of their models.

