[tooling] · · 2 min read
Amazon Is Buying and Destroying Rare Books to Feed Its AI Training Pipeline
A new report reveals Amazon is purchasing rare books, cutting off their spines, and scanning them, underscoring the growing demand for clean, human-written text.
By ByteBulletin Editors · Editorial Team
Amazon, the company that started as an online bookstore, is reportedly buying up rare books, slicing off their spines, and feeding them through scanners to train its AI models. The investigation, led by 404 Media, used a tracking device hidden inside a rare book to follow its journey to Amazon's Las Vegas facility, VGT3 — a site marked with a dinosaur clutching a book.
Amazon's statement to 404 Media was characteristically vague: it “purchases books through commercial channels to improve the products and services customers use.” But the implication is clear: the e-commerce giant is hungry for training data that hasn't already been consumed by the internet's AI gristmill.
Why Rare Books?
Large language models need vast amounts of text, and much of the open web has already been scraped. Publishers and authors have sued AI companies over copyright infringement, and Anthropic is currently facing litigation over pirated books. Rare, out-of-print books offer a legal — or at least a gray-market — source of text that is not only scarce but also valuable for a critical reason: it's pre-2022 human writing.
Text produced before the LLM boom is a safeguard against what researchers call “model collapse.” When models train on AI-generated content, their outputs can degrade, becoming repetitive and less useful. Rare books, with their unique style and vocabulary, are a potential antidote to homogenized data.
The Book-Destruction Problem
Cutting the spines off books to make them easier to scan is common in digitization projects, but at scale, it means the physical destruction of rare and potentially irreplaceable works. While libraries have long digitized books — often with or without permission — the commercial motivation here is different. This is not preservation; it's raw material extraction.
Book collectors and preservationists are alarmed. “It's one thing to digitize a book, but to physically destroy a rare book for data is a loss to cultural heritage,” said one expert, who asked not to be named.
For developers and technologists, this story is a reminder that the AI supply chain has real-world costs. The next frontier of training data isn't just the web or pirated ebooks — it's physical objects, shipped through Amazon's own logistics network. The irony is hard to miss: the company that disrupted bookstores is now disrupting books themselves.
SHARE
RELATED

[tooling] ·
UL-SMF: A Hardware-Software Fabric That Squeezes KV Cache Down to 2.6% of Its Size
A new open-source memory compression fabric claims up to 384x KV-cache reduction with >94% semantic retention, aiming to unblock long-context inference.
[tooling] ·
Browser-Native Image Tools Are Quietly Killing the Upload-and-Wait Converter
WebAssembly ports of MozJPEG, libwebp and libavif now run at near-native speed inside the browser tab — and a new generation of image tools is using them to keep your photos off other people's servers entirely.

[tooling] ·
Google overhauls hacker codenames: say goodbye to APT1, hello to 'Castle' and 'Ion'
Google's revamped naming system for hacking groups aims to bring clarity to a crowded field of threat actors.
