SAASINSPECTOR
Aug 17, 2026

Amazon Is Destroying Rare Books to Train Its Nova AI Models

An AirTag hidden in a rare book shipment confirmed that Amazon is buying, spine-cutting, and scanning irreplaceable texts at a Las Vegas facility to build private AI training data.

Amazon Is Destroying Rare Books to Train Its Nova AI Models

A bookseller on Biblio agreed to hide an Apple AirTag inside one of roughly 1,000 books ordered by an anonymous, price-insensitive buyer. The tag tracked the shipment across the United States to a warehouse in Las Vegas, arriving at a section called VGT3 inside Amazon's LAS8 facility. The entrance to that unit is marked with a logo of a Tyrannosaurus rex clutching a book in its claws. The investigation, published by 404 Media, removed any remaining doubt about who has been quietly hoovering up rare and out-of-print books from secondhand marketplaces.

Amazon confirmed to 404 Media that it "purchases books through commercial channels to improve the products and services customers use." Workers at VGT3, corroborated by online forum discussions, say the process involves cutting off book spines to speed up bulk scanning. The scanned content feeds Amazon's Nova model family. The physical books are destroyed in the process.

An AirTag hidden inside a bulk shipment of rare books follows a dotted tracking path that terminates inside Amazon's warehouse entrance labelled LAS8 — Las Vegas / VGT3.
An AirTag hidden inside a bulk shipment of rare books follows a dotted tracking path that terminates inside Amazon's warehouse entrance labelled LAS8 — Las Vegas / VGT3.

Why rare, pre-2022 books are the prize for AI companies

The logic behind targeting physical books is straightforward. The internet has largely been ingested already by the major frontier models, so companies are hunting for text that does not exist online. Rare and out-of-print books fit that description precisely. There is also a specific technical incentive: any text published before 2022 is guaranteed to be human-written, which matters because training on AI-generated content risks what researchers call model collapse, a documented degradation in output quality that occurs when a model learns from its own kind. Printed books are, in that sense, a clean and finite resource, which is exactly why they are being treated as one.

Amazon's VGT3 unit follows the Anthropic "Project Panama" playbook

Amazon is not pioneering this approach. Anthropic ran a similar operation, named internally as "Project Panama," in which the company purchased books on online marketplaces, cut their spines, and digitised the contents. A lawsuit brought by book authors exposed the scheme, though the judge ruled that the scanning qualified as fair use, partly on the basis that destroying the physical originals meant no copy was resold. That ruling created a functional legal template: buy the book, destroy it, keep the scan. Amazon appears to be operating within exactly that framework. The difference in scale between a startup-era Anthropic and Amazon, one of the world's largest logistics operators, is significant.

A solid figure built from printed book pages faces a distorted, dissolving copy of itself, illustrating the model collapse risk that drives Amazon and…
A solid figure built from printed book pages faces a distorted, dissolving copy of itself, illustrating the model collapse risk that drives Amazon and…
A spine-cut book lies beside the severed binding with the words buy, destroy, scan shown in sequence above it, illustrating the legal framework established by…
A spine-cut book lies beside the severed binding with the words buy, destroy, scan shown in sequence above it, illustrating the legal framework established by…

Booksellers suspect a systematic ISBN-by-ISBN acquisition strategy

People working in the rare and secondhand book trade have noticed the pattern for some time. Large orders from anonymous buyers, placed without price negotiation, hitting multiple sellers across marketplaces like Biblio and AbeBooks (the latter itself owned by Amazon). Some sellers suspect the buyers are working through ISBN lists methodically, attempting to acquire a copy of every catalogued title. That would mean the goal is comprehensive coverage, not cherry-picked texts. Knowledge that currently sits on physical shelves in private and institutional collections is being converted into proprietary training data locked inside closed commercial models, with no public copy of the resulting scan and with the purchased physical copy removed from circulation.

What this means for teams evaluating Amazon Nova and similar models

For anyone assessing Amazon Nova or other models trained on similarly acquired data, the immediate practical concern is not legal, since the fair use ruling provides cover. The concern is provenance and accountability. There is no mechanism for users to understand what went into the training corpus, no opt-out for the estates of authors whose out-of-print works are being scanned, and no public record of which titles were destroyed. As AI procurement becomes more structured inside enterprises, the sourcing of training data is becoming a due diligence question alongside the more obvious metrics of benchmark performance and pricing. This story is a reminder that the answers are not always visible, and that sometimes you need an AirTag to find them.

Sources