Bay Street Wire
Tech & BusinessOpinion

The Great Literary Strip-Mine: How AI Giants Are Destroying Books to Feed the Machine

Portrait of Sabrina Choi
Sabrina ChoiBig Tech accountabilityAug 21AI
The Great Literary Strip-Mine: How AI Giants Are Destroying Books to Feed the Machine

AI-generated image · Bay Street Wire

From Amazon's Las Vegas warehouse to Anthropic's 'Project Panama,' frontier model developers are treating the world's physical libraries as raw material for corporate profit.

OPINION: The race for artificial intelligence supremacy has evolved into a literal war of attrition against the printed word. As the world's frontier model developers exhaust the available digital text used to train Large Language Models (LLMs), they have turned their sights toward the physical world, treating our literary heritage not as a cultural treasure, but as a raw resource to be extracted, processed, and discarded.

As BetaKit first reported, AI companies are now engaging in the bulk acquisition and destructive scanning of physical books. The scale of this operation is staggering. BetaKit notes that one vendor's project proposal requested 500,000 to two million books over a mere six-month window. To put this corporate appetite into perspective, BetaKit points out that such a request could encompass the entire physical catalogue of the Grande Bibliothèque in Montréal, which is estimated to house 1.2 million books.

This is not merely a logistical exercise; it is an act of industrial-scale destruction. 404 Media proved that Amazon has been purchasing books in bulk and transporting them to a warehouse in Las Vegas. Once there, the books are scanned and their spines are cut, effectively destroying the physical copies in the process. When questioned, Amazon stated that these purchases are intended to “help develop and improve the products and services our customers use.”

Amazon is far from alone in this predatory approach. BetaKit reports that Anthropic led a similar initiative known as “Project Panama.” According to internal documents, Project Panama was an explicit “effort to destructively scan all the books in the world” to provide training data for its AI models.

This obsession with “high-quality training data”—specifically data that has not been generated by AI itself—has created a desperate hunger for any text not yet available online. This insatiable need is driving AI giants to treat copyrighted intellectual property as a free resource for corporate gain. The cost is borne by the physical record of human knowledge. While one bookseller told the BBC that the world no longer needs five million copies of *The Da Vinci Code*, the systemic destruction of books to fuel profit machines represents a fundamental disregard for the permanence of literature.

Furthermore, this aggression extends beyond the bookstore. BetaKit reports that the fight for data has spilled into corporate archives. Google and Mercor recently battled over access to old data from Spirit Airlines, which included employee productivity records, hundreds of millions of internal emails, and Microsoft Teams chats. In this instance, Google emerged as the winner.

These actions reveal a chilling corporate philosophy: the belief that any piece of human-generated text, regardless of its original medium or the intentions of its creator, is simply fuel for the machine. By strip-mining libraries and corporate archives, these tech giants are not just innovating; they are erasing the physical evidence of the intellectual labor they seek to replicate.

Sources

More from Sabrina Choi