← The Vault
The Big Story

Why AI companies are turning to physical books for data

Amazon has been purchasing rare, out-of-print books and destroying them to scan their pages for AI training data. This process happens because tech companies are running out of high-quality human writing on the internet. By using older books, companies can avoid "model collapse," a phenomenon where AI systems become less intelligent after training on too much content generated by other AI models. It is a strange irony that the bookstore giant is now dismantling history to feed its algorithms.

Edition № 425Room: The Big Story18 August 20262 min readSources: 1
Article

Amazon began its life as an online bookstore, but today it is dismantling rare, physical books to feed its artificial intelligence. Recent reports confirmed that the company has been purchasing books, removing their bindings, and scanning their pages to provide data for its AI systems.

WHAT'S HAPPENING

Amazon is acquiring rare and out-of-print books to use as training material for its large language models — the complex systems that power tools like ChatGPT or Amazon's own AI assistants. To turn these physical objects into digital data, the books are systematically destroyed by having their spines cut off so the pages can be fed through high-speed scanners. This process transforms a physical archive into a mountain of digital text that the AI can then process, analyze, and learn from. The company has acknowledged that it acquires books through commercial channels to improve its digital products.

Why AI needs to read old books

HOW IT WORKS

Training an AI is essentially a process of teaching it how human language works by showing it massive amounts of examples. The goal is to give the AI enough patterns to recognize so it can eventually predict the next word in a sentence accurately. So far, these companies have relied on the internet for this information, but they have largely run out of high-quality, human-written text. They are now hunting for fresh, reliable data sources that exist outside of the web. Older, rare books are highly prized because they were written long before AI existed. This is crucial because if an AI is trained on content created by another AI, its performance can suffer. This is known as model collapse, where the quality of an AI's output degrades over time, similar to making a photocopy of a photocopy until the image becomes blurry and unrecognizable. By using books published before the era of AI-generated content, companies hope to keep their systems sharp, accurate, and truly human-like.

WHY IT MATTERS

This practice highlights a strange tension in the tech industry: as AI companies struggle to find new data to stay competitive, they are turning to the very physical artifacts that modern technology once rendered obsolete. For a company that built its fortune on selling books to now be destroying them to fuel a future of automation, it raises a quiet question about what we value more: preserving the physical history of our written word or using those pages as raw fuel for the next generation of software. As the supply of human-generated information online dries up, we should expect to see more of our physical world converted into fuel for these digital machines.

Sources
← PreviousWhy Google is bringing in outside AI experts for ChromeNext →Why a major AI hardware startup is changing its business
Tomorrow's edition · free

Liked this one? The next lands at breakfast.

Every story in tomorrow's AI news, rebuilt in plain English — five minutes, sources linked, free forever.

By joining you agree to receive Article's daily newsletter — unsubscribe in one click. Privacy

← Back to the Vault