Washington | 23°C (clear sky)
AI Firms Are Buying and Shredding Rare Books to Train Their Models

Antique volumes are disappearing as AI labs scan, ingest, and then destroy them en masse

A surge in AI‑training pipelines has tech companies buying cheap copies of old, out‑of‑print books, ripping the pages, digitizing the text, and discarding the originals—raising copyright, preservation, and ethical concerns.

For most of us a book’s value comes from the stories it holds or the way it looks when it sits on a bookshelf. To the engineers building giant language models, however, a book is simply a massive data dump—something to be devoured, digitized, and then tossed aside.

It isn’t a brand‑new, glossy paperback that a startup pulls from an online store. Instead, many AI labs have turned to the back‑room of the used‑book market, snapping up piles of cheap, often out‑of‑print volumes. The most notorious example is Anthropic, which, after a $1.5 billion settlement over similar practices with digital texts, started buying physical copies, feeding them into a hydraulic cutter that ripped out the pages, and then scanning the sheets with industrial‑grade cameras. In legal terms the company leaned on the “first‑sale doctrine,” arguing that once a buyer owns a book they can do what they like with it, and a judge agreed the process was “transformative” enough to count as fair use.

What makes this whole operation scale is a little‑known service called ISBNdb. Originally built to help libraries and independent shops locate titles, the platform now markets itself as a bulk‑order hub for AI firms, promising deliveries ranging from a thousand to a million books per shipment. As one representative admitted, the optics are terrible—"AI company destroys two million books" doesn’t exactly win public sympathy—so ISBNdb keeps the buyers’ identities under wraps.

Booksellers are feeling the ripple. One small vendor recounted how his weekly sales jumped from a couple of dozen titles to hundreds almost overnight, the orders marked with nothing but ISBNs and an oddly random selection of old, foreign‑language works. He confessed to mixed feelings: the cash flow is welcome, but the thought of rare, perhaps unique copies being pulped is unsettling.

Across Europe, especially in the Netherlands, rare‑book dealers report a flood of bulk purchases that smell a lot like AI‑lab orders, even if they can’t prove it. The pattern is the same—massive, anonymous buying sprees for books published before the AI boom, supposedly because they’re “free of modern contamination.” The phrase “pre‑LLM era” has become a marketing buzzword, suggesting that texts written before 2022 are pristine training material, untouched by the very algorithms that now scour the internet.

The fallout is more than a legal grey area. Every scanned page is a lost artifact, and for some titles the original physical copy might have been the only surviving example. While digitization can be a boon for preservation, doing it en masse for profit—then discarding the source material—creates a new kind of cultural erasure. Critics argue that the industry needs clearer guidelines that balance the appetite for data with the responsibility to protect literary heritage.

In short, the AI boom is turning once‑cherished books into data points, and the quiet bulk‑order services that make this possible are only beginning to raise eyebrows. Whether lawmakers, publishers, or the tech giants themselves will step in remains to be seen.

Comments 0
Please login to post a comment. Login
No approved comments yet.

Editorial note: Nishadil may use AI assistance for news drafting and formatting. Readers can report issues from this page, and material corrections are reviewed under our editorial standards.