Something odd is happening in the secondhand book trade. Recently, there have been reports of AI companies buying secondhand books in large quantities, raising questions about their motives.
Booksellers across the UK and Ireland say they have been receiving unusually large orders for books that seem to have almost nothing in common. An obscure magazine issue might appear beside a literary translation and a decades-old specialist title. No obvious collector theme. No neat subject category.
And now some sellers are wondering whether the customers are really readers at all.
The suspicion: AI companies buying secondhand books as fresh training data.
The theory remains just that in many of these individual cases. The mystery buyers have not been publicly identified as AI companies. Still, the timing — and what has recently emerged about how major AI developers obtain books — has made the possibility harder to dismiss.
Booksellers Are Seeing Orders That Don’t Look Normal
Barter Books in Alnwick, Northumberland, began noticing unusual orders roughly three months ago, according to The Guardian. Co-owner Stuart Manley said normal bulk purchases often share a theme. These didn’t.
One recent order jumped between an Estonian translation of a John le Carré novel, a particular edition of Anne Brontë’s Agnes Grey and an October 1983 issue of Warship magazine. Barter Books has reportedly sold hundreds of books to three such buyers, bringing in roughly £4,000. It’s not an isolated story.
MW Books in Claregalway, Ireland, has seen large and apparently unrelated orders since May. BookLovers of Bath reported a burst of demand between May and July involving around 200 books. Another UK bookseller told The Guardian that buyers associated with these unusual orders had purchased about 6,000 books since January. The customers aren’t necessarily behaving like bargain-hunting book dealers either.
According to one seller, buyers have been willing to pay full prices even when purchasing substantial quantities. Different customer names have also reportedly sent books to the same delivery location, including a postcode linked to freight warehouses near Heathrow Airport. That’s where the story starts getting stranger.
Why AI Is Suddenly the Obvious Suspect
AI models need enormous quantities of text. The internet supplied plenty of it. But the internet of 2026 is no longer the internet of 2016.
Web pages, social posts, product descriptions, search results and even entire websites increasingly contain AI-generated material. For developers trying to train the next generation of language models, that creates an awkward problem: AI systems risk learning from content produced by other AI systems.
Old physical books don’t have that problem. A novel printed in 1987 cannot contain ChatGPT-generated prose. Neither can a technical manual from 1974 or a forgotten history book sitting in a secondhand warehouse. That makes dusty shelves unexpectedly valuable.
A July report from 404 Media highlighted this emerging market, describing how ISBNdb had promoted physical books as particularly attractive AI training material because they contain edited, pre-generative-AI human writing. The company’s pitch targeted high-volume book acquisition for AI-related data use, although ISBNdb later told The Guardian that the webpage involved was a test of market interest and had been removed.
Suddenly, a random stack of old books isn’t just inventory. It’s a dataset.
Anthropic Already Showed How Far AI Companies Will Go for Books
There is another reason booksellers are connecting these orders with artificial intelligence: Anthropic. Court records and reporting published earlier this year revealed an operation inside the Claude developer known as Project Panama.
Anthropic purchased millions of physical books. The books were then processed so their pages could be scanned into digital form for AI development. Documents described industrial equipment capable of cutting bindings before high-speed scanning, with processed books eventually sent for recycling. The scale was enormous.
The Washington Post reported on Anthropic’s book-scanning operation, including spending of tens of millions of dollars acquiring books, while one proposed scanning operation envisioned processing hundreds of thousands to millions of books within months.
Anthropic has said Claude is trained using a mixture of publicly available web information, commercially acquired datasets and data produced by the company itself. The company has also said acquiring books is an established approach within large-language-model development and that its programs do not purchase and destroy rare or antiquarian books.
That history doesn’t prove Anthropic — or any particular AI company — is behind the current mystery orders. It does make the basic idea very believable.
The Used Book Market May Have Accidentally Become Part of the AI Data Race
For years, discussions about AI training data mostly revolved around websites, journalism, social media, artwork and copyrighted material scraped from the internet. Physical books change the picture.
Huge numbers of titles remain poorly digitized or completely absent from easily accessible online datasets. Some contain specialist material that may be difficult to find anywhere else. Others were published decades before today’s web existed. Then there’s another advantage: human provenance.
Before the generative AI boom, publishers, editors and writers produced these books without worrying about whether the text itself had been synthesized by a language model. In a digital ecosystem increasingly filled with machine-generated material, that distinction could become commercially valuable.
It’s a strange inversion. AI was supposed to make physical information feel obsolete. Now the growing presence of AI-generated information online may be making physical books useful again. Not necessarily because anyone wants to read them. Because machines might.
Booksellers Are Getting the Money — But Some Aren’t Comfortable With the Destination
For a secondhand bookseller, a customer arriving and buying hundreds of slow-moving titles sounds like good news. Financially, it often is. But some sellers are uncomfortable with the possibility that books they have preserved for years could be stripped apart, scanned once and recycled.
Others take a more practical view: somebody purchased the book legally, and what happens afterward is the buyer’s decision. There’s another complication. Nobody has publicly demonstrated that every strange order now appearing across secondhand marketplaces is connected to AI. Some probably aren’t.
The orders could involve resellers, archival projects, libraries, collectors or other commercial operations. The unusual patterns have created suspicion, not proof. Still, the volume and randomness are enough that sellers on different continents have noticed similar behavior. Reports have emerged from the UK, Ireland, Australia, Europe and the United States.
AI’s Hunger for Human Data Isn’t Going Away
The bigger story isn’t really about a few thousand books moving through online marketplaces. It’s about what happens when the easiest training data has already been collected.
AI developers want better models. Better models need better data. And useful, reliably human-created material is becoming harder to separate from the growing ocean of synthetic content online. That pushes the search somewhere else. Private datasets. Licensed archives. Scientific material. Specialist collections.
And apparently, perhaps, the shelves of secondhand bookshops. For booksellers watching baffling orders arrive from unfamiliar buyers, the AI boom may have entered their businesses through an unexpected door. One obscure paperback at a time.

