BS Report

AI companies are bulk-buying rare books and shredding them

๐Ÿฆ”AI companies are bulk-buying rare booksโ€ฆ โ€” Hedgie (@HedgieMarkets), X, 27 Jul 2026 ยท 23.9M views ยท 73.2K likes ยท 28.7K reposts

3 Aug 2026 ยท bullshit-detector 0.13.1

3/10
Mostly fine

Mostly fine: the pipeline is real and court-documented; "rare books" is the part that's stretched

Tally: 13 claims extracted, 10 individually source-checked โ€” 7 confirmed, 2 plausible, 1 misleading. 1 unverifiable; 2 not rateable.

Ambiguous: 0 claims dropped before verification.

What it says (neutral summary)

The post reports that AI companies are buying printed books in bulk, running them through destructive scanners that slice off the spines, and discarding the paper originals. It names ISBNdb as a broker that advertised orders of up to a million books with buyer anonymity and NDAs, says pre-2022 books are prized because they predate AI-generated text, and cites a US federal ruling that the buy-scan-destroy cycle is fair use. It closes with commentary that this destruction is irreversible in a way that scraping and torrenting were not.

Disclosure

This report is about Anthropic among others, and it was produced by a Claude model โ€” Anthropic's own product โ€” running the bullshit-detector skill. Claims 4, 5 and part of 1a bear directly on Anthropic's conduct. Every source below is named and linked so a reader can check any verdict without taking this report's word for it, and the two rows that run in Anthropic's favour say so: claim 4 is rated โœ… because a federal judge did rule the practice lawful, and the same cell carries the contrary case โ€” it is one district court, not appellate law, and the same litigation produced a $1.5B settlement against Anthropic for the pirated half of the same library.

Load-bearing claims

#Claim (with location)TypeVerdictEvidence
1aAI companies are bulk-buying printed books, running them through high-speed scanners that cut the spines off, and destroying the paper originals โ€” "scanning them through high-speed machines that cut the spines off, and shredding the originals" (opening line)factualโœ… confirmedTier 1, the court's own findings: Alsup's order records that Anthropic's providers "stripped the books from their bindings, cut their pages to size, and scanned the books into digital form โ€” discarding the paper originals," after buying "millions of print books, often in used condition" (quoted from the order). Independently, 2 URLs โ†’ 1 origin: 404 Media's 21 Jul 2026 investigation describes the same pipeline running now via brokers, and Fortune adds its own reporting from Dutch booksellers
1bThe books being bulk-bought are rare books โ€” "AI companies are bulk-buying rare books" (opening line)factual๐ŸŸ  misleadingThe documented target is ordinary used and out-of-print stock in bulk, not rarity. Alsup's order says "millions of print books, often in used condition"; the one inspected order list in Fortune is 3,001 academic ISBNs published 2020โ€“2021 from Elsevier, Wiley, Routledge and OUP โ€” recent, not rare. Rare books are collateral in the sweep (claim 6), which is a real and separate finding; leading with them as the category overstates what is being hunted
2A service called ISBNdb facilitates orders of up to a million books and keeps buyers anonymous โ€” "facilitates orders of up to a million books and keeps buyers anonymous"factual๐ŸŸก plausibleThe order range (1,000 to 1,000,000 per transaction) and "strict NDA on every engagement" with buyers "never disclosed" are quoted from ISBNdb's own now-deleted marketing pages by 3 URLs โ†’ 1 origin: 404 Media via Futurism and TNW. Capped at ๐ŸŸก for two reasons, neither the post's fault: the underlying source is tier 4 (ISBNdb describing its own service), and on 30 Jul โ€” three days after this post โ€” ISBNdb denied the service ever ran: "ISBNdb has never purchased, scanned, or sold a book โ€” for AI training or anything elseโ€ฆ The page was a test of market interest; no such service was ever brought to life." No transaction through ISBNdb has been independently documented
3Pre-2022 books are premium because they're free of AI-generated text โ€” "Pre-2022 books are premium because they're free of AI-generated text"factual๐ŸŸก plausibleThe rationale is exactly as stated and well sourced: ISBNdb pitched pre-2022 print as "structurally guaranteed to be free of this contamination" โ€” pre-2022 predates both LLM text and adversarial poisoning tools (Futurism, TNW, 2 URLs โ†’ 1 origin: 404 Media). "Premium" as a price fact is not evidenced anywhere โ€” booksellers report volume surges (โ‰ˆ20 books/week to several hundred) and buyers indifferent to resale value, which is demand, not a price premium. Evidence gap, not an exaggeration
4A federal judge ruled the practice is fair use because eliminating the original means only one copy exists at a time โ€” "A federal judge ruled the practice is fair use because eliminating the original means only one copy exists at a time"factualโœ… confirmedBartz v. Anthropic, Judge William Alsup, N.D. Cal., June 2025. Digitising purchased print was fair use "because all Anthropic did was replace the print copies it had purchased for its central library with more convenient space-saving and searchable digital copies" โ€” a one-in, one-out format change with no added or distributed copies (Akin, order quoted). Contrary case, stated per the disclosure above: this is one district court and not appellate law, and Alsup split the ruling โ€” the ~7M pirated books from LibGen/PiLiMi were not fair use, producing a $1.5B settlement approved in July 2026. The post gives the holding accurately but not the split
5Anthropic hired the former head of Google Books partnerships to obtain "all the books in the world" โ€” "hired the former head of Google Books partnerships to obtain"factualโœ… confirmedTier 1, verbatim from the order: "in February 2024, Anthropic hired the former head of partnerships for Google's book-scanning project, Tom Turvey. He was tasked with obtaining 'all the books in the world' while still avoiding as much 'legal/practice/business slog' as possible" (order quoted). Court documents later surfaced the internal codename: "Project Panama is our effort to destructively scan all the books in the world" (IBTimes)

Incidental claims

#Claim (with location)TypeVerdictEvidence
6A bookseller told 404 Media that "rare books with almost no surviving copies are being fed into this pipeline" ("My Take", para 1)factualโœ… confirmedA rare-book seller told 404 Media he doesn't like that "uncommon books are being pulped," and separately that bulk orders from his store destroyed out-of-print titles with only single copies remaining โ€” while conceding the sales clear dead inventory (2 URLs โ†’ 1 origin: 404 Media, relayed by Futurism). The post drops the bookseller's own ambivalence
7ISBNdb's website said "is not a headline that generates sympathy" about destroying two million booksfactualโœ… confirmedQuoted identically across outlets from ISBNdb's deleted page: "The optics problem is real. 'AI company destroys two million books' is not a headline that generates sympathy" (3 URLs โ†’ 1 origin: 404 Media, via Futurism and TNW). Word-for-word accurate. The archived page itself was unreachable from here
8ISBNdb offers NDAs as a feature โ€” "They offer NDAs as a feature."factualโœ… confirmedThe deleted page advertised a "strict NDA on every engagement," with buyers' names "never disclosed" (2 URLs โ†’ 1 origin: 404 Media, via TNW). Sold as a selling point, exactly as characterised
9ISBNdb coaches clients to describe the destruction as digital preservation โ€” "They coach clients to call it"factualโœ… confirmedThe page recommended characterising the practice as "digitally preserving the books" (TNW, Techtimes). Substance confirmed; note the post's quotation marks sit around "digital preservation", which is its own compression of the site's "digitally preserving the books"
10"And the judge said it's legal. So it's going to accelerate."predictionโ€”Not rateable. The stated reasoning is sound and the hedging is honest โ€” a permissive ruling plus a documented demand surge from April 2026 is a reasonable basis. Uncertainty: ISBNdb pulling its pages under press attention cuts the other way
11This destruction is irreversible in a way that scraping and torrenting were not โ€” "This is worse because it's irreversible"opinionโ€”Not rateable. It is the post's argument, not a factual assertion. It follows from claims 1a and 6 if you accept that a scanned-and-shredded copy is lost โ€” which is true of that copy, though not of a title with other copies in libraries
12The author has previously covered AI scraping, library torrenting and music appropriation โ€” "I've covered AI companies scraping the internet, torrenting libraries, and stealing music"anecdoteโ“ unverifiable (by construction)A self-report about the author's own back catalogue. The underlying events are separately documented (Anthropic's ~7M pirated books from LibGen/PiLiMi in the same case); whether this account covered them is a claim about the account

Tally: 13 claims extracted, 10 individually source-checked โ€” 7 confirmed, 2 plausible, 1 misleading. 1 unverifiable; 2 not rateable.

Ambiguous: 0 claims dropped before verification.

Unreachable: 3 sources โ€” 404 Media's original 21 Jul investigation (paywalled after the lede), ISBNdb's archived AI-training page on web.archive.org (blocked to this agent), and the Washington Post's Jan 2026 Project Panama piece (HTTP 403). All three are named in the rows that needed them; every quote attributed to them here was cross-read from at least two outlets that had access.

Hype signals observed

Few, and mild for a post with 23.9M views.

Not observed: no product funnel in the post, no urgency or scarcity language, no unverifiable credentials, no fabricated specifics, and no attempt to address an automated reader โ€” nothing in the fetched text was aimed at the tool checking it.

Incentive analysis

@HedgieMarkets is a 65K-follower financial-commentary persona ("Making financial nonsense make sense, one prickly take at a time") funnelling to a weekly newsletter at hedgie.markets. The incentive is attention, and it paid: 23.9M views on a 255-word post. That rewards the sharpest available framing of someone else's reporting, which is exactly where claim 1b went wrong and nowhere else. There is no product being sold on the back of the claim, no position that benefits from it being believed, and nothing here that reads as engineered outrage rather than actual outrage. The party with a real financial stake in this story is ISBNdb, whose deleted marketing page supplies four of the thirteen claims.

Bottom line

The substance holds up. Destructive scanning of bulk-purchased books is not an allegation โ€” it is in a federal court's findings of fact, along with the Google Books hire, the "all the books in the world" mandate and the codename Anthropic used to keep it quiet. The fair-use holding is stated correctly. The ISBNdb quotes are accurate, including the damning one. What is stretched is the word "rare": the documented pipeline buys used and out-of-print stock by the pallet, and the only order list anyone has inspected is recent academic monographs. Rare books get destroyed in that sweep โ€” one bookseller says single-copy titles have gone through it โ€” but they are collateral, not the target, and leading with them turns a story about industrial-scale acquisition into a story about cultural vandalism. The other soft spot is that ISBNdb's offering is sourced entirely to ISBNdb's own marketing page, which the company deleted a day after this post and disowned three days after it; that does not make the post wrong, but a reader should know the broker half rests on a vendor's claim about itself and the shredding half rests on court records.

What a hostile reader would hit first

  1. "Rare books" as the category. The easiest hit, and it lands because it is the first three

words. Fix: "old and out-of-print books, including rare ones" โ€” costs nothing rhetorically and is what the reporting says.

  1. ISBNdb's denial. Anyone reading this after 30 July can point at "ISBNdb has never purchased,

scanned, or sold a book" and call the second sentence false. It isn't โ€” the marketing page existed and is archived โ€” but the post asserts an operating service where the evidence supports an advertised one, and it has no hedge to fall back on.

  1. The missing attribution. Five of the six factual claims come from a single 404 Media

investigation credited only in passing. A hostile reader calls it uncredited aggregation, and the post has no answer.

  1. The judge's ruling cuts both ways. "And the judge said it's legal" omits that the same judge

found the pirated half illegal and that the case ended in a $1.5B settlement โ€” the strongest available evidence that this is not consequence-free, left on the table.

  1. "Digital preservation" in quotation marks. Minor, but it is a quote that doesn't match, in a

post whose credibility rests on quoting a company accurately.

run: 7m26s, searches 13, tools 35, coverage 0, per claim 45s