Book spines sliced, pages scanned, originals discarded at Amazon’s Las Vegas VGT3 unit
By Roger Satterfield
TECH TIMES

Amazon.com founder and CEO Jeffrey P. Bezos speaks at an event unveiling the new Amazon Kindle 2.0 at the Morgan Library & Museum February 9, 2009 in New York City. Mario Tama/Getty Images
Amazon once built its entire business selling books. It now runs a warehouse in Las Vegas where rare ones are sliced apart and fed to its AI models — and a reporter proved it with a $29 tracker. A 404 Media investigation published Monday tracked rare books to Amazon by placing an Apple AirTag inside a bulk order of approximately 1,000 volumes suspected of heading to an AI company, following the shipment across four states to an Amazon facility called LAS8 in northeast Las Vegas, where a dedicated unit identified by the internal code VGT3 receives pallet shipments, cuts the spines off books to speed scanning, and discards the physical copies. The scanned pages are used to train Amazon’s AI models. The books do not survive.
The finding is significant for a specific reason: Amazon had previously denied engaging in destructive scanning. The AirTag contradicted that denial directly.
What the AirTag Found at Amazon’s LAS8
The tracking device’s journey was unglamorous and damning. A rare-book seller on Biblio — an independent marketplace that keeps buyer identities anonymous — cooperated with 404 Media reporter Emanuel Maiberg to slip the AirTag into one volume in a 1,000-book bulk order that the seller suspected was destined for AI training.
The device flew first to Milwaukee, then sat for two weeks in a distribution warehouse outside Kenosha, Wisconsin. A truck carried it west with an overnight stop in Grand Junction, Colorado. Its final signal came from Las Vegas, from the north end of an Amazon warehouse complex called LAS8, as GeekWire reported.
LAS8 is primarily a print-on-demand facility — a place that creates books. But the north end of the building runs a separate operation. Amazon employees who posted on an internal workers’ forum identified the unit by its code name: VGT3. The logo painted on VGT3’s entrance is a Tyrannosaurus rex holding an open book. The image is accurate advertising.
“All we do is scan books,” one employee wrote on the forum, as reported by Tom’s Hardware. Other workers described the workflow: large pallet shipments of books arrive, staff slice the spines off each volume so the loose pages can feed through high-speed scanners, and the physical copies are discarded once the scan is complete. Nothing comes out the other end except a digital file.
When both 404 Media and Ars Technica asked Amazon for comment, the company provided the same statement to each: “Amazon purchases books through commercial channels to help develop and improve the products and services our customers use,” according to The Next Web. The statement does not mention artificial intelligence. Nothing in it contradicts the reporting.
What the Barcode Scan Reveals
One detail in the 404 Media reporting does more work than any other.
VGT3 employees do not simply scan the books’ pages. According to worker accounts reviewed by 404 Media, staff scan each book’s ISBN barcode before scanning its pages. The International Standard Book Number — a 13-digit identifier assigned to every commercially published book since 1970, standardized as ISO 2108 and now covering more than 150 countries — is the global cataloging system that makes every published book uniquely identifiable.
By scanning ISBNs before content, Amazon is building a record of what it has consumed. That record, cross-referenced against any ISBN database, tells the operation which books it still needs. An AI company with access to an ISBN catalog can, in principle, generate systematic purchase orders to fill every gap in its training corpus — working through the entire documented bibliography of human publishing one pallet at a time. Booksellers across the US and Europe have already described exactly this pattern: bulk orders with no coherent subject matter, unified only by having ISBNs, placed by buyers with no interest in the books’ literary or commercial value, as documented in TechTimes’ earlier coverage.
The bookseller who planted the tracker described what gets lost in this calculus. “There are different types of value,” the seller told 404 Media, as reported by The Next Web. “There’s monetary value, obviously, but there are a lot of other types of value. There’s historical value, intellectual value, sentimental value. All sorts of things, and all of those the AI companies don’t care about. They just want the content as a bunch of words strung together.”
Why the Open Web No Longer Works
To understand why Amazon is running a book-shredding operation in the Nevada desert, you need to understand what has happened to the internet as a training resource.
The first generation of large language models was built primarily on web crawls. That resource is now compromised in two distinct ways. The first is model collapse: a documented phenomenon in which AI models that train on AI-generated text degrade progressively across generations, because the synthetic output of earlier models lacks the distributional diversity of genuinely human-authored text. Researchers at Oxford, Cambridge, Imperial College London, and the University of Toronto published the formal analysis in Nature in July 2024, warning that the feedback loop was already self-reinforcing. An Ahrefs study of approximately 900,000 newly created web pages found that 74.2% contained AI-generated content as of April 2025 — meaning the contamination had already spread to nearly three-quarters of what any new web crawl would find.
The second problem is adversarial poisoning. Techniques pioneered by the University of Chicago’s SAND Lab — the image-domain tool Nightshade demonstrated the general principle — have enabled authors to embed invisible corruptions in their text that degrade AI model training in targeted ways. A joint study by Anthropic, the UK AI Security Institute, and the Alan Turing Institute found that as few as 250 malicious documents in a training corpus of trillions of tokens can plant a measurable backdoor.
Pre-2022 printed books sidestep both problems simultaneously. They were written before any large language model existed, meaning they cannot contain AI-generated text. They were also written before any data-poisoning tool existed, meaning they cannot have been adversarially corrupted. Researchers at Epoch AI have projected that demand for public human-generated text could exhaust the available supply sometime between 2026 and 2032. Books printed before the AI era represent a finite, fixed, and irreplaceable corpus that is structurally clean in ways no contemporary web crawl can match.
ISBNdb — the metadata firm that briefly offered bulk book-acquisition services to AI companies before 404 Media’s earlier investigation prompted a retraction in July 2026 — was candid about this logic in its now-deleted marketing materials: “The world’s best AI training data is sitting on a shelf.” The firm noted that pre-2022 print books were “structurally clean” precisely because they predate both contamination sources, as covered in TechTimes’ ISBNdb investigation.
How Anthropic Did It First — and What Amazon Has Refused to Say
Amazon did not invent this practice. It is following a path that Anthropic blazed, documented in federal court.
Internal documents from Anthropic unsealed in January 2026, as part of the copyright class-action lawsuit Bartz v. Anthropic, described an internal initiative called Project Panama in a single extraordinary line: “Project Panama is our effort to destructively scan all the books in the world,” as detailed by TechTimes. Anthropic hired Tom Turvey — the former head of partnerships for the Google Books project — to lead the effort, deliberately replicating the legal framework that allowed Google’s book-scanning to survive copyright challenges. The mechanism was identical to what Amazon’s VGT3 uses: hydraulic cutting machines remove the spines; high-speed scanners process the loose pages; the physical books are recycled.
In June 2025, U.S. District Judge William Alsup ruled in Bartz v. Anthropic that converting a lawfully purchased physical book into a digital AI training file constituted transformative fair use under copyright law — a court-endorsed green light for destruction-based acquisition. The Alsup ruling is a single district-court decision from the Northern District of California; because Anthropic settled the case rather than appealing it, the ruling never reached an appellate court and creates no binding precedent for other jurisdictions. The U.S. Copyright Office expressed skepticism about fair use in a May 2025 report that not all commercial AI training qualifies. But for AI companies operating now, it functions as provisional legal cover.
The Bartz settlement was finalized in July 2026, with Anthropic paying approximately $1.5 billion — $3,000 for each of roughly 500,000 covered works — making it the largest copyright class-action in US history. That payout covered the piracy part of the case: Anthropic separately downloaded more than seven million books from pirate repositories like Library Genesis and Pirate Library Mirror, which the court found was not protected by fair use.
The critical difference between Anthropic and Amazon is not legal. It is a single statement. When 404 Media and the BBC asked Anthropic about its practices, the company specified that its data acquisition programs do not buy and destroy rare or antiquarian books. xAI — Elon Musk’s AI company — has publicly said the same.
Amazon has made no such statement. Its spokesperson’s comment — “purchases books through commercial channels to improve the products and services customers use” — does not specify what categories of books are purchased, what physical condition they arrive in, what happens to the physical copies, or what AI application receives the scanned content. The bookseller who placed the AirTag sourced the books from Biblio, a marketplace that specializes in rare and used titles.
Read more: Anthropic Copyright Settlement Gets Final Approval: $3,000 Per Book, No Binding Precedent
Is Amazon Destroying Books Legally? The Short Answer Is Yes — For Now
The legal framework that makes VGT3’s operation possible is the same one that protected Anthropic’s Project Panama: a combination of the first-sale doctrine (a buyer may dispose of a physical copy however they wish) and the Alsup ruling (converting a purchased physical book into a digital AI training file is transformative fair use).
Amazon’s operation, as reported, appears to stay within that framework: it purchases books through commercial channels — it does not appear to be downloading pirated copies — and it destroys the physical originals while creating a single digital version. The Alsup ruling specifically noted that destroying the physical copy while scanning was consistent with the one-copy-to-one-copy logic of fair use.
That legal picture may not hold internationally. Emily Hudson, an intellectual property specialist at the University of Oxford, told the BBC that UK law requires copyright permission for both creating the training dataset and conducting the training — no fair use equivalent exists. German copyright law is even clearer: Germany’s Publishers and Booksellers’ Association has stated that scanning books would violate German copyright whether or not the physical books were legally purchased.
For the authors whose work arrives in Amazon’s Las Vegas facility, the legal clarity does not resolve the ethical gap. An author whose book was lawfully purchased, sold on a secondary market, and then shredded in Nevada receives no notification and no compensation. The Bartz settlement covered Anthropic’s piracy; it did not address what fair use authorizes. Authors who sold or licensed their works to publishers and retained rights over secondary uses have no mechanism to opt out of a lawful bulk purchase from a used-book marketplace.
What Happens to Books That Are the Only Copy Left
The AirTag investigation did not track a copy of a bestseller with a million surviving editions. It tracked a shipment sourced from Biblio — a marketplace that specializes in titles that are difficult to find elsewhere.
The distinction matters. For a book with 500,000 existing copies, losing one purchased copy to a scanner is trivially recoverable. For a specialist, scholarly, or regional publication with 100 surviving copies — or fewer — the calculus is different. Once a rare book is spine-cut and shredded, the text may survive as a private data corpus. The physical object, and everything that distinguishes a physical book from a string of tokens — the marginalia, the binding, the provenance, the material history — is gone permanently.
Booksellers in Europe have already identified the specific pattern that makes rare-book targeting detectable: buyers who pay above market price for low-monetary-value titles, who request books unified only by having ISBNs, and who decline to identify themselves or their intended use. Derek Walker of McNaughtan’s bookshop in Edinburgh told The Next Web he had sold what appeared to be the only known surviving copy of an 18th-century edition through a channel he now suspected was AI procurement — and that this would be a substantially different problem than selling common stock.
Stuart Manley of Barter Books in Northumberland offered the trade’s most candid counterpoint. “The world no longer needs five million copies of The Da Vinci Code,” he said, noting that he had titles that sat unsold online for 20 years and had finally moved through bulk orders. The industry is not united in opposition. It is genuinely divided between sellers who see a welcome market for slow-moving stock and those who recognize that rare and specialist books are a categorically different matter.
Amazon has drawn no line between the two categories.
What Comes Next for Amazon’s AI Strategy — and for This Legal Landscape
Amazon’s timing on the VGT3 operation is complicated by developments in its AI strategy. As of late July 2026, Amazon began winding down its flagship Nova models — including Nova Premier, Nova Omni, Nova Reel, and Nova Canvas — redirecting resources toward a single new frontier model effort led by researcher Pieter Abbeel, expected to debut at Amazon’s re:Invent conference later this year. Active Nova models include Nova 2 Lite, Nova 2 Sonic, and Nova Forge. Amazon has confirmed it continues to invest in model development; the scanned books would feed whatever model emerges from the Frontier Model Research program.
The legal landscape surrounding physical book acquisition is also in motion. The Bartz ruling established that scanning is fair use but created no binding precedent. A new class-action filed in New York federal court on July 10, 2026 — by Hachette Book Group, Cengage Learning, Elsevier, and author Scott Turow against Google over Gemini training data — is not bound by California precedent. That complaint includes allegations that Google removed or altered copyright information on ingested works.
For used-book sellers, the practical consequence of Monday’s investigation is a name. Booksellers have spent a year noticing the pattern — incoherent bulk orders, anonymous buyers, prices that made no commercial sense — without being able to identify who was placing them. One AirTag changed that. The destination was an Amazon warehouse in Las Vegas. The unit was VGT3. The operation is ongoing.
Frequently Asked Questions
Is it legal for Amazon to buy rare books and destroy them for AI training?
Under current US law, yes — with important caveats. A June 2025 ruling by U.S. District Judge William Alsup held that converting a lawfully purchased physical book into a digital AI training file constitutes transformative fair use under 17 U.S.C. § 107, as detailed in TechTimes’ settlement coverage. This ruling is a single district-court decision from California that has never been reviewed by an appeals court and creates no binding precedent for other courts. The U.S. Copyright Office has expressed skepticism that all commercial AI training qualifies as fair use. Outside the US, the legal picture is significantly different: UK and German law do not contain equivalent fair use exceptions for AI training, meaning the same operation conducted in Europe would likely be illegal.
Why are AI companies specifically buying books published before 2022?
Pre-2022 books offer something the open web no longer can: a guarantee of purely human authorship. Researchers documented in Nature in 2024 that AI models trained on AI-generated content degrade across generations — a phenomenon called model collapse — and an Ahrefs study found that 74.2% of newly created web pages contained AI-generated text as of April 2025. Pre-2022 books also predate adversarial data-poisoning tools, which can embed invisible corruptions into training data. Physical books printed before the AI era are structurally clean in ways no contemporary web content can be — which makes them increasingly valuable as a training resource and increasingly targeted for acquisition.
What did Amazon previously say about destructive scanning, and how does the AirTag change that?
Amazon had previously denied engaging in destructive scanning, according to reporting by UniladTech. The company’s current statement — “purchases books through commercial channels to help develop and improve the products and services our customers use” — does not mention AI, does not describe the process, and does not specify what happens to the physical books. Amazon employees described in internal forum posts that VGT3’s entire function is cutting book spines and scanning the loose pages, with the physical books discarded afterward. The AirTag confirmed those accounts by physically tracking a shipment to the facility. Amazon has offered no explanation for the discrepancy between its prior denial and the operation documented by the investigation.
What can used-book sellers do if they suspect an order is for AI training?
No law currently requires buyers to disclose their identity or intended use in a standard commercial book purchase. However, sellers who are concerned can look for the characteristic pattern that booksellers across the US and Europe have described: bulk orders with no coherent subject matter unified only by having ISBNs; buyers who show no price sensitivity and will not negotiate; anonymous purchasing identities; and shipping destinations inconsistent with the books’ stated value or subject. Booksellers who decline such orders — as Tomás Kenny of Kennys Bookshop in Galway has done — are exercising the only form of veto available in a supply chain specifically designed to obscure its purpose. Whether to fill such an order is ultimately a business and ethical decision each seller must make without legal guidance, in the absence of any regulatory body that has moved to address this market.