Somewhere in the secondhand-book trade, the economics of an old book have changed as buyers pursue titles that conventional demand largely ignored. Booksellers in Britain and Ireland have reported large, eclectic orders paid at full price and sometimes routed through freight warehouses. One British seller said buyers fitting the pattern had ordered roughly 6,000 books since January 2026. Their ultimate use is not established in every case, but sellers increasingly suspect that some demand is tied to artificial-intelligence data acquisition.
That suspicion has a documented precedent in litigation involving Anthropic, where a federal court described millions of printed books being purchased, stripped of their bindings, scanned and discarded. The court treated those lawfully purchased books differently from a separate library of pirated copies, making acquisition method central to the copyright dispute. The case does not settle the legality of AI training broadly, but it shows why physical books can become useful when desired material is absent from convenient digital channels.
Australian booksellers have reported similar patterns and a narrower preservation risk. Destructive scanning can save the words while eliminating physical evidence carried by a particular copy, including annotations or provenance. There is no evidence that AI companies systematically target rare books, but scarce material can lose historical context when the object itself is destroyed.
The unsettling image of a book surviving for decades only to lose its spine gives the economic shift a physical form. Yet the books are the warning, not the point. They reveal a new consumer assigning value to knowledge by different rules.
| Acquisition Signal | Documented Scale |
|---|---|
| Anthropic print purchases | Millions of books |
| Print acquisition spending | Many millions of dollars |
| Pirated digital library | More than 7 million copies |
| Reported UK seller orders | About 6,000 books |
Sources: United States District Court Northern District of California, The Guardian
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
When a Machine Values a Book Differently
Traditional secondhand markets price books around human interest, with readership and collectability often determining whether an old title retains demand. Artificial intelligence introduces another valuation question: does the book contain useful information that cannot be obtained easily elsewhere?
That question can reverse the economics of obscurity because low readership does not imply low informational value. A technical manual may preserve terminology that never migrated to the searchable web, while a local history may record facts barely represented in digital archives. Machine demand can therefore reward informational gaps that human markets largely ignored.
The unusual purchasing patterns become more coherent under that logic, even though no dataset yet establishes a market-wide AI premium for obscure books. A biography and an agricultural treatise need not share a subject if both fill missing areas in an existing corpus. Informational uniqueness can acquire value independently of the audience for the object containing it.
Older books may also carry provenance advantages because they predate the mass spread of generative AI across the information environment. That does not make older material inherently superior, and current evidence does not show a broad premium for pre-AI books. It does mean origin becomes more consequential as synthetic text grows easier to produce and harder to distinguish from human material.
Artificial intelligence can create text at immense scale while increasing the relative scarcity of differentiated information with credible provenance or specialist authority. By lowering the cost of finding and applying neglected knowledge, machines can make obscurity economically relevant rather than commercially fatal.
| Measure | Scale |
|---|---|
| Languages spoken worldwide | 7,000+ |
| Languages represented online | About 1,000 |
| Web content concentrated in 14 languages | More than 91% |
| People still offline | 2.6 billion |
Sources: UNESCO, ICANN
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
The Return of the Information Silo
For much of the internet era, information trapped inside an institution or private database was treated as an inefficiency to overcome through digitization and search. Artificial intelligence changes part of that logic because useful knowledge can remain external to a model while still shaping its answers.
Retrieval-augmented generation allows an AI application to search an external collection and provide relevant material to a language model as grounding context. Microsoft describes the approach as a way to use proprietary or non-public information without embedding every useful document inside the model itself.
Economically, external retrieval separates ownership from use and creates room for controlled access rather than permanent transfer. An institution can sell a copy of a dataset, or it can keep the collection and allow machines to consult it under defined conditions. The second model preserves more control over the underlying knowledge.
Possession alone, however, does not create a valuable AI asset. An archive must become reliable enough for machines to use: scans need accurate text, documents need context, rights need clarity, and retrieval needs to work. UK guidance on AI-ready government data similarly treats ownership and provenance as foundations for machine use rather than administrative details.
A warehouse containing 20,000 anatomy books remains a library until the collection becomes searchable, trustworthy, legally usable and difficult to reproduce. At that point its economic character begins to change because value lies in controlled access to organized knowledge, not merely in possession of documents.
| Europeana Measure | Scale |
|---|---|
| Digital heritage items | 59 million+ |
| Data providers | 3,500+ |
| Content freely reusable | 42% |
| Records raised to quality standard in one year | 3 million |
Sources: Europeana
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
The Next AI Advantage May Be What a Model Can Reach
Generative-AI competition has centered heavily on model performance and the computing resources required to produce it, but retrieval systems create another potential source of differentiation. Two models with similar general capabilities can produce different results when one can consult authoritative information unavailable to the other.
That asymmetry changes incentives for developers and information owners. Developers can acquire or license differentiated collections, while owners can reconsider whether transferring a permanent copy captures the full value of what they control. Firms may build defensible AI services around specialized knowledge without competing directly to build frontier models.

The relevant scarcity is often not the information itself because digital documents can be copied cheaply without depleting the original. Scarcity can instead reside in lawful machine-use rights, authenticated provenance or continuing access to a corpus that competitors cannot readily reproduce.
Copyright disputes reinforce the commercial importance of those distinctions even while broader training law remains unsettled. Anthropic’s litigation separated model training from the acquisition and retention of pirated books, making lawful provenance economically significant when information becomes an AI input.
A mature global market for privileged information access has not yet formed, and specialist knowledge does not replace model quality or computing power. As capable models spread, however, trusted information can become a stronger source of differentiation. Competitive advantage may increasingly depend on what a model is permitted to reach.
| Data Arrangement | Reported Value | Information Type |
|---|---|---|
| Google and Reddit | About $60 million a year | User discussions |
| OpenAI and News Corp | Reported at $250+ million over 5 years | Current and archived news |
| Shutterstock Big Tech deals | Initially $25–50 million each | Images, video and music |
| Defined.ai long-form video | $100–300 per hour | Licensed training media |
Sources: Reuters
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Who Owns the Knowledge AI Has Not Yet Found
Institutions outside the frontier-model race may possess information poorly represented in current AI systems, including local knowledge or specialized research developed outside dominant digital channels. Material that sits at the edge of one information economy can become strategically useful in another.
Many institutions still treat such collections mainly as preservation responsibilities or administrative costs rather than as potential components of AI strategy. UK policy already separates possession from machine readiness by tying useful datasets to clear ownership, licensing and provenance.
The resulting policy problem resembles natural-resource economics only in a limited sense because information is not depleted when it is copied. Control can still shift when rights are transferred before owners understand future demand. Bargaining power may disappear even while the original documents remain physically or digitally intact.
Institutions therefore face choices about the terms of machine access rather than a simple choice between secrecy and unrestricted release. They can preserve originals while digitizing collections, distinguish public reading from commercial machine use, and license access without surrendering permanent possession.
For organizations sitting on decades of overlooked material, the emerging economic question is what becomes more valuable when machines begin searching systematically for what human markets neglected.
| U.S. National Archives Measure | Scale |
|---|---|
| Estimated textual pages | 12.25 billion |
| Textual scans online | 462 million |
| Estimated textual holdings digitized | 3.78% |
| Electronic records held | 33+ billion |
Sources: U.S. National Archives and Records Administration
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
The Books Are the Warning
The strange secondhand-book market gives that question a physical form because a low-value object can suddenly enter a different economic system. A book survives for decades, reaches a dealer with little human demand and then attracts a buyer interested primarily in whether its information exists elsewhere.
When the binding comes off and the pages become data, the object may disappear while its informational content acquires a new use. Most information silos will not announce that transition so visibly, since a scientific archive or corporate knowledge base can gain machine value without leaving its owner.
The broader shift is that information once difficult to find, organize or monetize can become more useful when machines reduce the cost of locating and applying it. The first AI race concentrated on systems capable of learning from enormous quantities of available information; a growing part of the next may concern access to what was never readily available.
Books disappearing into scanners are therefore more than a story about digitization or destruction. They show artificial intelligence beginning to alter the economic value attached to possessing human knowledge.
| Value Dimension | Useful Metric | Higher Value Signal |
|---|---|---|
| Information scarcity | Share unavailable online | Large offline share |
| Provenance | Sources with verified origin | High verification rate |
| Rights clarity | Corpus with defined machine-use rights | High rights coverage |
| Machine readiness | Searchable, structured records | High usable share |
| Replication difficulty | Comparable external sources | Few substitutes |
| Access control | Permissioned retrieval share | Granular control |
Sources: UK Government Digital Service, Microsoft Azure
TL;DR Summary
- Secondhand booksellers have reported unusual bulk purchases of obscure and apparently unrelated titles.
- One British seller reported roughly 6,000 books ordered by buyers fitting the pattern since January 2026.
- Anthropic’s litigation established a documented precedent for purchasing, destructively scanning and discarding printed books for AI development.
- The evidence does not show that every unusual book order is connected to AI or that rare books are being systematically targeted.
- AI can assign value to informational uniqueness even when conventional human demand for a book is weak.
- Older books may carry additional provenance value because they predate the mass spread of generative AI content.
- Retrieval-augmented generation allows models to use external proprietary information without permanently embedding every document in model parameters.
- Controlled information silos can therefore become economically useful when their contents are reliable, searchable and legally usable.
- Scarcity may increasingly lie in machine-use rights, provenance and controlled access rather than in the ability to copy digital information.
- Firms controlling specialized knowledge may gain an AI advantage without competing directly to build frontier models.
- Governments, libraries and universities may possess overlooked information whose future machine value exceeds its current market value.
- Books disappearing into scanners provide an early physical signal that AI is changing the economics of possessing human knowledge.
Sources
- The Guardian; Secondhand booksellers in UK and Ireland suspect AI firms behind strange bulk orders; – Link
- The Guardian; More than just objects Australian booksellers raise alarm over destruction of rare titles to feed AI; – Link
- U.S. District Court Northern District of California; Bartz et al v Anthropic PBC Order on Fair Use; – Link
When a Machine Values a Book Differently
- British Library; 50 Facts About the British Library; – Link
- British Library; Our Collections; – Link
- Institute of Internet Economics; AI and the End of the Webpage; – Link
The Return of the Information Silo
- Microsoft Azure; Retrieval Augmented Generation in Azure AI Search; – Link
- UK Government Digital Service and Department for Science Innovation and Technology; Guidelines and Best Practices for Making Government Datasets Ready for AI; – Link
- U.S. National Archives and Records Administration; Record Group Explorer; – Link
The Next AI Advantage May Be What a Model Can Reach
- Reddit; Quarterly SEC Filing Revenue Contract Disclosures; – Link
- Reuters; Inside Big Tech’s Underground Race to Buy AI Training Data; – Link
- OpenAI and News Corp; A Landmark Multi Year Global Partnership with News Corp; – Link
Who Owns the Knowledge AI Has Not Yet Found
- OECD; Enhancing Access to and Sharing of Data in the Age of Artificial Intelligence; – Link
- OECD; Digital Government Index and Open Useful and Reusable Data Index 2025 Results and Key Findings; – Link
- UK Department for Science Innovation and Technology; National Data Library Progress Update January 2026; – Link
The Books Are the Warning
- UNESCO; Digitization Projects; – Link
- U.S. Copyright Office; Copyright and Artificial Intelligence; – Link
- U.S. District Court Northern District of California; Bartz et al v Anthropic PBC Final Approval of Class Action Settlement; – Link
- Institute of Internet Economics; Authors Rally Against AI’s Unauthorized Use of Millions of Books; – Link