“AI Companies are Buying […] Old Books” – and apparently then destroying them
Another topic I did not want to get involved in, other than locking my own fics to users only. Alas, here we are and I don’t even know why I am surprised.
To break it down, it looks like AI companies have been buying up tons of old books to dismantle, scan and then destroy them. And apparently, they’ve been buying from book resellers and second-hand bookstores, which means that some of those books they’ve been destroying may well have been some of the last copies outside of a library or archive.
Below is a collection of articles talking about this topic. There are paywalls on some of them, but I trust in the broader internet to find a way to access them nonetheless:
Secondhand booksellers believe they may have been caught up in the AI supply chain that sees old books scanned then destroyed
ISBNdb, a company that sources printed books for AI companies to turn into training data, tells clients “the optics problem is real.”
AI Companies Are Destroying Books. The “Antique Book” Panic Is Another Story.
Another day, another headline about artificial intelligence seemingly designed to trigger immediate outrage:
“AI Companies Are Buying Antique Books, Ingesting Their Contents to Train Models, and Then Destroying Them at Incredible Scale.”
That sounds horrifying.
It also goes considerably further than the available evidence.
There is a real story here. At least one AI company has purchased and destroyed millions of physical books after scanning them. Booksellers have also reported unusual bulk orders that they suspect may be connected to AI training.
But “AI companies are destroying books” and “AI companies are destroying antique books at incredible scale” are not the same claim.
Let us separate what has been documented from what has been assumed.
What actually happened
In 2025, court documents revealed that Anthropic had spent millions of dollars purchasing millions of physical books, often second-hand.
The company’s contractors removed the bindings, cut the pages to a uniform size, scanned them, and discarded the paper copies. Each physical book was replaced with a searchable digital copy in Anthropic’s internal research library.
That part is true.
Anthropic really did buy and destroy millions of books. This is not speculation based on unusual orders or anonymous sources. It is described directly in a federal court ruling.
However, the ruling dealt with several legally distinct actions.
The court found that using books to train Anthropic’s language models was fair use under the circumstances presented. It also found that replacing each legally purchased print copy with one internal digital copy was fair use because the original was destroyed, the total number of copies did not increase, and the digital replacement was not distributed outside the company.
The court did not find that Anthropic was entitled to download and permanently retain millions of pirated books merely because some might eventually be used for AI training. Those piracy claims were later covered by a US$1.5-billion settlement that received final approval in July 2026.
So, yes: lawfully purchased books were destructively scanned.
But no: the ruling did not give AI companies a universal licence to acquire books however they please.
Where did the “antique books” claim come from?
The recent panic began with reports about unusually large and apparently random book orders.
One bookseller told 404 Media that his sales had suddenly risen from fewer than twenty books a week to hundreds. The orders showed little apparent interest in subject, author, genre, or even reasonable pricing. The books reportedly had ISBNs, leading the seller to suspect that they were being purchased in bulk for some form of automated processing.
The seller also carried foreign-language, uncommon, and out-of-print books. He therefore worried that some of the copies might be scarce and could be destroyed after scanning.
That concern is reasonable.
It is not proof that antique or irreplaceable books are being destroyed at an “incredible scale.”
Even Futurism’s article acknowledges that booksellers generally do not know who placed the orders or what happened to the books afterwards. The buyers may be AI companies, but the sellers are largely inferring that from the unusual purchasing patterns.
Some uncommon books may be disappearing into private scanning operations.
That possibility deserves investigation.
It should not be reported as a proven worldwide assault on antiquarian collections.
An ISBN is not a rarity certificate
The fact that the reported bulk orders focused on ISBN-listed books is also important.
An ISBN is a commercial product identifier. It identifies a particular title, edition, publisher, and format so that booksellers, libraries, distributors, and other participants in the book trade can catalogue and order it.
It does not tell us whether an individual copy is rare, valuable, historically important, or one of thousands gathering dust in second-hand shops.
ISBNs were issued as ten-digit numbers until the end of 2006 and as thirteen-digit numbers from 2007 onward. Older ten-digit ISBNs can also be converted into thirteen-digit equivalents, so a modern database displaying an ISBN-13 does not necessarily mean that the book itself was published after 2007.
A book can have an ISBN and still be scarce or out of print.
It can also be an extremely common paperback that nobody has wanted since 1993.
“ISBN-listed,” “out of print,” “rare,” and “antique” are not interchangeable descriptions.
The alleged middleman has now walked back its claims
Much of the recent reporting focused on ISBNdb, a book-metadata company whose website advertised bulk book-acquisition services for AI developers.
Its marketing pages claimed it could source between 1,000 and one million books per order. They promoted books published before the recent proliferation of generative AI as a supply of supposedly uncontaminated human-written material. They also discussed confidentiality and acknowledged that headlines about AI companies destroying millions of books would create an “optics problem.”
Those pages certainly existed.
But after 404 Media reported on them, ISBNdb removed the material and issued a denial. The company now says the pages were only a test of potential market interest, that the service was never launched, and that it has never purchased, scanned, or sold a physical book for AI training.
That denial does not prove that no other company is sourcing books for AI laboratories. It does mean that we should not cite ISBNdb’s abandoned marketing pitch as proof that such purchases are already occurring through that company at the advertised scale.
A company advertised a possible service.
Booksellers noticed strange orders.
Some sellers suspect AI buyers.
Those facts justify questions.
They do not justify presenting every suspicion as a confirmed industrial operation.
Book destruction did not begin with AI
There is another piece of context missing from much of the outrage: the book trade already destroys books routinely.
Bookstores commonly return unsold stock to publishers or wholesalers. Some returned copies are resold or remaindered. Others are pulped, recycled, or otherwise destroyed because storing, shipping, inspecting, and redistributing them would cost more than they are likely to earn.
An industry report on international bookselling found that many returned books are pulped, while also noting how little transparency exists concerning the number destroyed each year. The International Publishers Association describes pulping unwanted titles as an ordinary outcome of the returns system.
The same thing happens in the second-hand economy.
Charity and thrift shops cannot keep every donated book indefinitely. Shelf and warehouse space are finite. The Salvation Army has openly explained that books unsuitable for resale may be sent to cardboard and paper-pulp recyclers.
None of this makes Anthropic’s destruction of books admirable.
It does expose a glaring double standard.
When publishers pulp unsold stock because they printed more copies than consumers wanted, it is treated as an obscure and unfortunate feature of the publishing business.
When an AI company buys used books, scans them, and destroys the originals, it becomes evidence that technology is devouring human culture.
The physical outcome may be the same.
The symbolism is different.
Why the symbolism matters
Books are not perceived as ordinary consumer goods, even when the industry treats them that way.
A warehouse full of unsold paperbacks quietly sent for recycling rarely produces viral headlines. An industrial machine cutting the spines from books so that an AI company can extract their contents is much more emotionally potent.
It looks like a machine literally consuming literature.
That reaction is understandable. Books can be cultural artefacts, personal objects, historical records, and vessels of human expression.
But the emotional power of an image does not relieve journalists of the obligation to report accurately.
If the concern is the destruction of culturally important books, then the relevant questions are whether the books are genuinely scarce, whether other copies have been preserved, and whether the resulting scans will remain accessible.
Simply calling them “antique” does not answer any of those questions.
Does buying a book make AI training legal?
Lawfully purchasing the books matters, but it is not a magic permission slip.
In the United States, one federal court ruled that Anthropic’s training use was fair use under the facts before it. The court separately ruled that replacing legally purchased print copies with one-for-one internal digital versions was fair use.
Those were related but distinct legal conclusions.
Other countries do not necessarily use the American doctrine of fair use. Some instead have specific exceptions for text and data mining.
European Union law permits text and data mining of lawfully accessible material under certain conditions, although rightsholders may reserve some uses. The United Kingdom currently permits text and data mining without additional permission only for non-commercial research where the researcher already has lawful access.
The accurate conclusion is therefore not:
Buying a book automatically makes commercial AI training legal everywhere.
It is:
Lawful acquisition can be an important part of whether computational analysis is permitted, but the answer depends on the jurisdiction, the purpose, and the exact use being made of the copies.
That may be less exciting than a sweeping declaration.
Copyright law has never been famous for its thrilling simplicity.
What deserves genuine scrutiny
There are still legitimate issues here.
Companies should be transparent about where their training material comes from and how it is acquired.
Books that are genuinely scarce, historically significant, uniquely annotated, or otherwise important should not casually disappear into inaccessible corporate archives.
Destructive scanning may also be wasteful when non-destructive methods are practical, particularly if the physical book would otherwise remain available to readers.
And a privately held scan is not the same thing as preservation. A book removed from circulation and converted into proprietary training data has not necessarily been preserved for the public.
Those concerns are serious enough without pretending that AI companies have already been proved to be systematically shredding the world’s antique treasures.
The truth does not need help from clickbait
The confirmed story is already strange and uncomfortable.
Anthropic legally purchased millions of physical books, destroyed them while making digital replacements, and used books from its larger library in developing language models.
Other booksellers have seen mysterious bulk orders that may be connected to similar projects.
Some uncommon books could be involved.
But there is currently no reliable evidence demonstrating that AI companies are destroying antique or irreplaceable books at an enormous global scale.
That conclusion was created by stretching a concern into a certainty, a suspicion into an industry-wide fact, and “out of print” into “antique.”
Criticism of AI companies should be based on what they actually do.
When legitimate concerns are inflated into clickbait, the result is not stronger accountability. It is weaker journalism — and another reason for readers to distrust the next alarming headline, even when that one might be true.
Sources
Bartz v. Anthropic PBC: federal court order on AI training, purchased print books, destructive scanning, and pirated digital libraries. [Link 1] [Link 2]
Reuters: final approval of the US$1.5-billion Anthropic copyright settlement. [Link]
404 Media: original ISBNdb report and the company’s subsequent denial and removal of its book-sourcing pages. [Link 1] [Link 2]
Futurism: the article claiming that antique books are being destroyed at an incredible scale. [Link]
International ISBN Agency: what an ISBN identifies and the transition from ten to thirteen digits. [Link 1] [Link 2]
International Publishers Association and RISE Bookselling: returns, recycling, pulping, and the lack of reliable industry-wide totals. [Link 1] [Link 2]
European Union and United Kingdom government guidance on text-and-data-mining exceptions. [Link 1] [Link 2]
Process pics!! My mom has a bunch of old cookbooks that I plan to slowly gut and turn into journals, or something. The microwave cooking ones were an obvious first choice.
I've watched some YouTube videos on how to take the pages out of a book, and there's kind of two ways that books are bound, so I removed the pages of the one that was bound the way that makes it the easiest to cut them out, and left the other one whole for now.
Then I covered them in gesso, and both of these started from just having a little leftover paint and not wanting to waste it, and then just seeing what happened from there. Also the mountain one had a lot from the cover that shined through the gesso so I wanted to make sure I did something dark over that and the shape reminded me of a mountain!
The waves are definitely still a work in progress... I'm having the most difficulty with understanding where the splashes should be. I think I need to practice with contrast and empty space, and restraint.
The mountain... I'm not sure! I'm trying to convey that the sun is rising behind the mountain, without it being too dark, but also without ignoring where the light source is actually coming from and showing differences because of that? I think I kinda got that on the trees and hills, but the flowers might be too bright... But also I love flowers so maybe that's just artistic talent expression. Like, flowers are so great they'll be bright when they wanna be, light source be damned.
The destruction of books is hardly a historical rarity. After the Spanish physician and theologian Michael Servetus was burnt at the stake in 1553 for religious heresy, only three surviving copies of his The Restoration of Christianity were spared from the bonfire. One copy, now held by the University of Edinburgh library, is ironically believed to have been personally saved by the Protestant leader John Calvin, who was responsible for Servetus’s execution. Calvin bragged in a letter that he had “purged the Church of so pernicious a monster.” He did not preserve the book out of affection, but rather due to a zealous, tyrannical bibliophile’s sentiment.
Yet, Calvin is still not as noxious as modern artificial intelligence companies, including Anthropic. These corporations engage in bulk purchasing, digitizing, and then pulping out-of-print books to train their large language models. Calvin at least had the excuse of his God; AI companies do not.
A sobering report by Frank Landymore in Futurism explains that Large Language Models (LLMs) are increasingly trained on their own "slop." This has created a desperate need to train them on pre-2022 writing that remains uncontaminated by AI itself. This is the exact sales pitch made by ISBNdb, a company supplying bulk orders of between one thousand and one million volumes at a time to tech giants.
Through a process of “destructive scanning,” a hydraulic cutting machine slices off book spines. The pages are scanned, digitized, and immediately pulped. This corporate laundering eludes copyright regulations through a legal loophole, meaning AI companies could actively be destroying some of the few remaining physical copies of these books.
While publishers already destroy nearly 25% of unsold works each year for purely financial reasons, this new revelation portends something far more terrifying: the rapacious, cannibalistic devouring of human culture to feed the very technology meant to replace us.
If a rare copy of Servetus’s book fell beneath the lifeless eye of the digitizer today, it would simply be uploaded into the hungry maw of the machine to serve as fodder for an advanced plagiarism device.
OpenAI CEO Sam Altman remarked, “We see a future where intelligence is a utility, and people buy it from us on a meter.” This is a technocratic war against the sacred materiality of the book, reducing history to predictive binary code.
Crucially, when physical books are systematically destroyed, the internet loses its final tether to absolute verifiability. There is no longer any objective proof of exactitude left on the net.
As a result, future generations will blindly trust digital data and online narratives rather than seeking out original, physical books they are no longer able to find or willing to read.
It is an act of literal erasure that restores George Orwell’s warning to its original impact: “who controls the past controls the future.” LLMs swell while the shelves starve, resulting in a profound libricide—a corporate indifference to the obliteration of culture, ideas, and humanity itself.
Picture: Unknown North Netherlandish artist, “Roundel with Allegorical Scene of Book Burning” (c. 1520–30), colorless glass, vitreous paint, and silver stain (image public domain CCO via the Metropolitan Museum of Art;
: ̗̀➛ disclaimer : I’m not ordering anyone around, I wrote this text with my heart of a passionate of literature, and it’s a reminder for myself first and foremost. I post it here also because my love for Tumblr grew over the past two months, I’ve felt comfortable enough to write here before any other platform - it could fit better on Substack for example but I feel too intimidated to post there.
I have deleted Instagram and Tiktok because I'm politically vocal about my opinions, yet sometimes what happens in the world overwhelms me and I need a break from the waves of negative news around - let's be honest, seeing the world inclining towards fascism at an alarming speed can drive everybody mad. Yet yesterday I've come back on Instagram for an hour, to send messages with my friends, and I've read the most terrifying news I've read since a while : AI companies are buying old books in bulk and destroying them to train their LLMs, and we lose then every trace of these books.
As I read this, my eyes went wide with horror : we can't even imagine how much we're losing right now by the time I'm writing, and who knows which piece of our history and civilization are being erased from the general public ? Because here's the thing : AI companies are so devious that they buy books no one knows, to which the production is most totally over. Secondhand bookstores have notice random rises in orders, to which rumors link to this monstrosity of a book-burning. Not even Ray Bradbury could predict that Fahrenheit 451 would be happening less than a century after the publication of his book.
So this post is my way to blow the whistle about this : I'm begging you to stop using any form of generative AI, become curious again and read books. Whether you ask ChatGPT even for the most obscure information you can't seem to find on the first five pages of your browser search, whether you use Character.AI to entertain yourself, whether you watch any form of AI slop on your platforms, whether you can't seem to find one fanart conveying the exact idea you have in mind, I'm begging you to stop entirely. Not only because of the disastrous ecological consequences, not either because of the racially-targeted areas where those data centres are built, but at least if you care about art, you should boycott those companies entirely.
I won't talk here about the concepts of ecology, climate change and climate justice surrounding this technology, many before me have done it so well. What my heart is heavy with is the fact that we're losing one of the purest and most ancient form of human expression : the arts and humanities. What testimonies we have from the past aren't only the new technologies such as irrigation systems, tools to make clothes, houses and architecture, but also books, religious parchments, drawings, paintings, statues, clothes, fashion, music compositions, instruments, everything that the human species created and that AI is actively destroying now. How little we know about ancient civilizations are pieces of art, no matter which one it is. I repeat it : art is the purest and most ancient forms of expression that the humanity has, and we're activeley destroying it.
To come back about the books - because it's the one art I know and love the most : become curious, get out of those AI machines and those algorithms and read more. It's not about how many books you can read, it's about how much. Read a little everyday, and don't stick to that echo chambers that algorithms can create for you. Read diversly, read old, very old books, from unknown authors, visit those antique stores, those vintages stores, those thrift stores, your local bookstore, your local library. Read classics, dark romance, fantasy, mangas, read that dusty book in the back of the shelves, bearing the name of your old neighbor as an author. Collect books, save them from that fate. There's no such thing as "book overconsumption". Every book, whether it's a shakespearean masterpiece or a "Booktok book" (I despise that term but you know what I mean) carries a part of the period it has been published. Every book is a direct witness of the time and ages it has been published, don't pay attention to the elitist voices telling you which one is worthy of your time and which is not, just read. We're in a time and age where distraction and attention are the top-one priorities of those greedy companies. Don't give them what they want, read, educate yourself, write, even if it's bad at first. During a period where reading has become a priviledge, we need to make it a normal thing to do.
Hoard these books, read them, exchange them with your sibling, your cousin, recommend them on your platform. Reading is one of the best tool we got now, don't let it go to waste. Don't let those companies orchestrate our very own Alexandria book burning. Don't let these dystopian authors predict our future. Read.