Return to site

Amazon Is Cutting Up Rare Books to Feed Its AI

August 17, 2026

A journalist slipped an AirTag into a rare book and followed it across the country. According to 404 Media's investigation, the tracker ended inside an Amazon facility where workers cut bindings off books so the pages scan faster, then destroy or discard the printed originals. The order covered close to 1,000 rare books purchased through the bookseller marketplace Biblio, and the shipment eventually reached Amazon's LAS8 warehouse and its VGT3 operation in Las Vegas [1].

Amazon told 404 Media that it "purchases books through commercial channels to help develop and improve the products and services our customers use," but did not answer follow-up questions about why the books are cut, how many facilities do similar work, or how the company responds to criticism of destructive scanning [1]. The investigation was subsequently covered by Ars Technica and TechCrunch, which both framed it as evidence that at least one major AI company is buying physical books to turn them into training data [2][3].

If you buy or oversee AI systems, this is not only a strange story about books. It is a preview of the next fight over training data.

What "rare" means here

The word "rare" can easily send the mind to museum cases and first editions. That is not what the reporting proves. 404 Media says it is not disclosing the titles, and describes the books as rare in the practical sense that not many copies are in circulation. Some may be obscure, out of print, written in less widely spoken languages, or hard to replace [1].

That distinction matters. There is no public evidence that Amazon destroyed any unique or final copy of a specific book. The stronger point is still serious: scarcity can give physical books historical, research, or community value that is not necessarily captured by a page scan. Once the spine is cut and the paper runs through a scanner, the original object is gone [1].

This is why the story landed so sharply. TechCrunch noted the irony that Amazon, a company that began as an online bookseller, is now reportedly cutting up books to train AI models [3]. Ars Technica picked up the same core finding from 404 Media: booksellers had suspected AI-related bulk orders, and the AirTag connected one of those shipments to Amazon [2].

Why old books are suddenly valuable to AI labs

Old printed books solve a data problem that the AI industry partly created for itself. The open web has been scraped heavily, and more of it now contains machine-generated text. Books published before the generative-AI boom offer a relatively clean source of human-written material, without the growing risk that scraped web pages were themselves generated by an LLM [3].

The Next Web, summarizing 404 Media's earlier reporting, described how ISBNdb markets old printed books to AI labs as dense, edited, authoritative material that is not already polluted by synthetic output [4]. For a model builder, that is attractive. For everyone else, it raises a harder question: what happens when the industry's hunger for clean data starts pulling scarce material out of the physical world?

Captain Taine might call this a case of taking soundings before sailing into shallow water. The legal chart is only one chart.

Legal is not the same as defensible

Companies have spent two years arguing about whether training AI on copyrighted material is fair use. That fight is still moving through courts. AP reported that a federal judge approved a $1.5 billion settlement with Anthropic over claims involving pirated copies of books, while an earlier ruling in the same case found that Anthropic's training use of lawfully acquired books qualified as fair use [5].

Amazon's situation differs in one important way: as far as the available reporting shows, it paid for these books through normal commercial channels. That is not piracy [1].

But "we paid for it" is a lower bar than most business buyers should accept. The next generation of AI due diligence will have to distinguish three questions:

  • Was the data legal to use?
  • Can the vendor trace where it came from?
  • Can the vendor defend how it was acquired?

Amazon's book-buying operation shows why those are no longer the same question. Legal clearance is necessary. It is no longer sufficient.

Amazon Nova is context, not proof

Amazon develops its own Nova family of foundation models, including the current Nova 2 generation, which AWS offers through services including Amazon Bedrock [6]. Nothing in the reporting proves this specific shipment fed a specific Nova model, and it would be a stretch to claim otherwise.

That caveat matters. The point is not to overclaim. The point is to notice what the reporting makes visible: large AI vendors need high-quality training data, and the sourcing path for that data may involve physical supply chains, third-party marketplaces, opaque vendor relationships, and choices the public never sees unless someone hides an AirTag in a box.

For procurement teams, the inability or unwillingness to answer specific provenance questions is itself information. It does not prove misconduct. It does create uncertainty that buyers have to evaluate.

If this post has you thinking less about AI features and more about AI oversight, the next useful skill is governance: how to ask better questions, set clearer policies, and review risk before a system becomes business-critical. The AI Governance course is built for leaders and professionals who want a practical frame for responsible AI adoption, accountability, and oversight.*

Four checks before you trust an AI vendor's data story

  1. Ask what categories of training data were used. Do not stop at "licensed" or "commercially acquired."
  2. Ask whether any source material was scarce, out of print, hard to replace, or physically destroyed during ingestion.
  3. Ask what provenance controls exist. Can the vendor trace source categories, acquisition standards, and treatment of physical source material?
  4. Treat vague answers as risk signals. If a vendor cannot explain its sourcing categories, acquisition standards, provenance controls, or treatment of physical source material, assume they either do not know or will not say, and price that uncertainty into the deal.

The AI data governance question is shifting from "Do you have the legal right to use this?" to "Can you defend how you acquired it?" That is a better question for buyers, boards, and anyone else trying to build AI without losing the plot.

Sources