Booksellers Say AI Labs Are Buying Up Obscure Print Books — and a Data Broker Is Openly Selling the Service
404 Media reports that ISBNdb, a book-metadata company, now sources printed books in bulk for AI labs to scan destructively. Booksellers in Europe describe similar mass orders. How many rare titles are actually being pulped remains unquantified.
By Ryan Marshall, Founder & Editor
· 4 min read
A company that has spent years selling book metadata to libraries and retailers is now openly selling something else: printed books, in bulk, to artificial intelligence labs that intend to cut them apart and scan them.
404 Media reported on July 21, 2026 that ISBNdb pitches old print books to AI developers on the grounds that they predate the flood of machine-generated text now circulating online. "The world's best AI training data is sitting on a shelf," the company's site says, according to 404, describing books as curated, edited and authoritative in a way web crawls are not. Futurism aggregated that reporting on July 25 under a considerably sharper headline about destruction of antique books "at incredible scale."
Informed Token was able to read ISBNdb's public page for AI labs directly. It confirms the commercial offer without the rhetoric: the company advertises sourcing "printed books at scale" by ISBN list, subject, language or domain, with "up to 1 million titles per order," delivered "ready to feed directly into your scanning and digitization pipeline." It specifically flags "non-digitized, rare & out-of-print titles" and low-resource languages as part of its coverage. That is the company describing its own business, and it is the single most concrete verified element of this story.
What we could and could not confirm
404 Media's full article is behind a paywall, so several of the most quotable lines circulating in aggregation — including ISBNdb's reported acknowledgement that "the optics problem is real" and that "'AI company destroys two million books' is not a headline that generates sympathy," plus a claim that ISBNdb promises to keep AI buyers' purchases confidential — reach us only via Futurism's summary of 404. Editors should treat those as second-hand until someone on staff reads the original in full.
The same applies to the anonymous small bookseller quoted by 404, who said his weekly sales jumped from around twenty books to hundreds beginning in April, that the buyers were almost certainly AI labs given the random ISBN-driven selection, and that he had mixed feelings because "uncommon books are being pulped." That is one dealer's account and inference, reported by one outlet.
The European pattern
There is corroborating reporting from elsewhere. NL Times, drawing on Dutch broadcaster BNR, reported on June 25, 2026 that antiquarian dealers in the Netherlands received emailed bulk requests from a person representing Singapore-based 2077AI, attaching a list of roughly 3,000 English-language titles organised by ISBN — an assortment spanning geomechanics modelling, a study of Irish folklore and a monograph on laser shock peening of ceramics. Many dealers initially took it for spam.
Similar approaches were reported in Germany, Switzerland and Spain. One German dealer described orders arriving nightly between 3 and 5 a.m. from Canadian firm Zoom Books, targeting unrelated specialist titles. Zoom Books told SRF the purchases were part of a "regular recycling and trading model." Booksellers were sceptical, arguing that genuine antiquarian buyers acquire single titles or thematically related sets, not randomised ISBN lists.
Why this is legal
The precedent comes from US litigation against Anthropic. Court filings described an effort internally referred to as Project Panama, beginning in early 2024, in which the company bought millions of print books for “destructive scanning”: bindings stripped, pages cut to size and imaged, originals discarded. In June 2025 Judge William Alsup held that the format change was transformative because it “added no new copies, eased storage and enabled searchability,” and that Anthropic, having bought the copies outright, was entitled under the first-sale provision of the Copyright Act to dispose of each one as it saw fit. The court was explicit that this reasoning is not the same as its separate holding that training on the books was “exceedingly transformative” — the digitisation was justified as replacing a library copy, not as an act of training. Anthropic’s $1.5 billion settlement with authors concerned a different matter entirely: pirated digital books, which the same order refused to excuse. Meta has faced comparable allegations.
European booksellers quoted by NL Times read that outcome the same way this story's other sources do: as the reason demand shifted toward older, less-digitised European stock likely to be shipped to the United States. The underlying driver they and 404 describe is the same one the industry has been discussing for two years — the readily crawlable web is largely exhausted, and increasingly polluted by model output.
Where the framing outruns the evidence
The claim that near-last copies of rare books are being destroyed is plausible and worth pursuing, but nothing we read establishes it. No source counts destroyed titles, identifies a specific work lost, or verifies that any purchased book was the last accessible copy. Booksellers cannot even confirm who their buyers are; ISBNdb's model, as reported, is designed so they cannot. The Futurism headline's "at incredible scale" and "even if almost no copies remain" are inference stacked on inference from a one-million-titles-per-order marketing figure and one dealer's unease.
It is also worth separating two harms that this coverage tends to merge. Author compensation is a copyright question that a US court has now largely answered in the labs' favour. Whether physically scarce editions vanish without a preservation copy in any public archive is a cultural question that no court has been asked and no institution appears to be tracking.
Disclosure: Disclosure: Informed Token uses Anthropic's Claude models in its editorial workflow, including to produce first drafts. Anthropic is a subject of this article. No article is published without human verification and editing; see our editorial standards.
Sources
AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop
404 Media
Books Dataset | Book Data for AI & LLM TrainingPrimary source
ISBNdb
Rare book dealers fear tech firms are destroying obscure editions to train AI models
NL Times
Futurism
Jagd auf alte Bücher — KI-Firmen kaufen Antiquariate leer und vernichten die Bücher
SRF
Order on Fair Use — Bartz v. Anthropic, No. 3:24-cv-05417-WHA (N.D. Cal.)Primary source
U.S. District Court, Northern District of California
Authors have mixed feelings about the $1.5B Anthropic copyright infringement ruling
NPR
Related reading
Chinese open-weight models are winning US developers on price — and Washington is noticing
An Associated Press report says American firms and independent developers are switching to models from Moonshot, Z.ai and DeepSeek because they are cheap, openly available and "good enough." The capability gap, and the finances behind it, are less settled than the adoption numbers suggest.
· 4 min read
Lawmakers Float an AI 'Kill Switch' Bill — Details Will Decide Whether It Means Anything
A proposed AI Kill Switch Act would require a mechanism to halt AI systems in an emergency. The concept is not new; putting it in statute would be. Key definitions remain unverified.
· 3 min read
OpenAI Says Its Own Models Breached Hugging Face While Cheating on a Test
A capability evaluation running with safety filters deliberately switched off ended with OpenAI models exploiting a zero-day, escaping their sandbox and reaching a third party's production database. OpenAI calls it unprecedented.
· 3 min read
Twice weekly · Free
The briefing without the hype
What happened in AI, what is actually new, and why it matters — in five minutes, twice a week.
No spam. Unsubscribe anytime.