AMAZON’S AI IS EATING RARE BOOKS

The story broke properly this week, and it is as ugly as it sounds. Amazon and other artificial intelligence outfits have been quietly buying up physical books by the truckload, many of them rare, out-of-print or simply hard to find, then slicing the spines off so the pages can be fed through industrial scanners and turned into training data for their language models. Once the machines have sucked the text out, the books themselves are discarded. Destroyed.

It is not rumour. A bookseller in the United States, fed up with the strange bulk orders that had been arriving for months, slipped an Apple AirTag into one volume among roughly a thousand rare titles and watched the package travel across the country. It ended at an Amazon facility on the edge of Las Vegas. Workers there have described the routine on internal forums: pallets of books arrive, barcodes are scanned, spines are cut, pages go through the scanners, and the paper is thrown away. The team that does this work uses a logo of a Tyrannosaurus rex clutching an open book. Amazon’s public response has been the usual corporate fog: they purchase books through commercial channels to improve products and services. They have not denied the destruction.

Anthropic, the company behind Claude, was already on the record. Court documents from a copyright case earlier this year showed an internal project called Panama whose stated aim was to “destructively scan all the books in the world.” They bought millions of volumes, preferred the less common ones, hired contractors to cut the bindings and feed the pages into high-speed machines, then shredded what remained. Company spokesmen later insisted they were not targeting true antiquarian rarities, only harder-to-find professional and out-of-print works. The distinction may comfort their lawyers more than anyone who has ever held an old book and felt its weight.

The motive is straightforward enough. The great language models have already vacuumed most of the open internet. What remains online is increasingly contaminated with text written by earlier versions of the same machines, and training on that material risks what researchers call model collapse. Books printed before the generative-AI boom offer clean human prose that cannot easily be scraped from websites. Buying the physical object, scanning it, and destroying it also creates a chain of possession that some lawyers believe is more defensible under American fair-use rules than simple piracy. A federal judge has already treated Anthropic’s practice as transformative in at least one ruling. The legal shield, if it holds, is bought at the price of the book itself.

Booksellers on both sides of the Atlantic and here in Australia have noticed the same pattern: sudden orders for obscure titles that had sat unsold for years, often in lots of hundreds or thousands. Some dealers are making money they never expected. Others are watching volumes disappear that may never be replaced. True one-of-a-kind rarities appear less often in these hauls than solid mid-tier out-of-print stock, yet the cumulative effect is the same. Knowledge that once existed only on paper is reduced to training tokens, and the paper is gone.

There is a non-destructive way to digitise a book. It is slower and more expensive. Speed and cost won. The companies involved are racing one another for the next leap in capability, and a warehouse full of sliced spines is simply another input cost. Amazon, which began as an online bookseller, now runs a facility whose daily work is the systematic destruction of the very objects that once defined it. The irony is almost too neat to be believed, yet the AirTag did not lie.

What is being lost is not merely paper. Every destroyed volume takes with it a particular arrangement of thought that existed in that form and no other. Digital copies can be searched and copied endlessly; they cannot replace the object that once sat on a shelf, carried marginal notes, or passed from one reader to another. When the supply of clean pre-AI text grows scarce, the temptation to treat every remaining physical book as raw material will only increase. The companies involved have shown they are willing to act on that temptation at industrial scale.

The rest of us are left with the ordinary questions that attend any act of cultural erasure. Who decides which books are expendable? What happens when the last accessible copy of a difficult or unfashionable text is fed into a scanner and then discarded? And how long before the same logic is applied to other repositories of the past that stand in the way of the next model release? The machines need more text. The books are convenient. The rest is paperwork. 