Amazon is buying rare and out-of-print books in bulk, using their text to train artificial intelligence systems, and then permanently deleting them. According to 404 Media, the company has systematized this practice as part of its AI development process. The books being targeted are often unavailable anywhere else, held in few physical copies or long out of circulation. Once Amazon deletes its digital version, the text vanishes entirely.
This practice matters because these books represent knowledge that becomes irretrievable. Researchers looking for a specific out-of-print title, students needing access to historical texts, and book collectors searching for rare editions all lose a potential source. Libraries traditionally served as custodians of such materials, buying rare books to preserve them for future generations. Amazon is doing the opposite: acquiring them temporarily, extracting value, and discarding them.
The story reveals how AI companies train their systems. These models require massive amounts of text data to learn language patterns and generate responses. Amazon, like OpenAI and Google, sources this data from books, articles, and web content. The difference is that while some of this sourcing happens transparently through partnerships, much of it does not. Books are purchased, their text is extracted, and they are deleted. Authors and publishers receive no compensation for their work being used to build commercial AI products.
This creates a contradiction at the heart of how AI is being developed. Companies argue they need access to published knowledge to build useful systems. Authors and publishers argue their work deserves protection and payment. The practical outcome is that companies like Amazon resolve this tension by simply erasing the evidence. The books cannot be used again by competitors or preserved for cultural purposes if they no longer exist.
For India and other countries developing AI industries, this practice raises policy questions. Should regulations require tech companies to preserve the texts they use for training? Should authors be compensated when their work trains commercial AI? Should there be mandatory deposits of training data with libraries or archives? These questions do not yet have answers in most countries, meaning companies operate without constraint.
What happens next depends on whether publishers, authors, or regulators push back. The book industry has already sued some AI companies for unauthorized use. Amazon’s practice of buying and destroying books may trigger similar action. Until then, rare and out-of-print works will continue disappearing into private AI systems and emerging nowhere else.

