Wednesday, September 9, 2026

Book Obsolescence and Destruction; Artificial Intelligence and the Economics of Information Silos

Must Read

Somewhere in the secondhand-book trade, the economics of an old book have changed as buyers pursue titles that conventional demand largely ignored. Booksellers in Britain and Ireland have reported large, eclectic orders paid at full price and sometimes routed through freight warehouses. One British seller said buyers fitting the pattern had ordered roughly 6,000 books since January 2026. Their ultimate use is not established in every case, but sellers increasingly suspect that some demand is tied to artificial-intelligence data acquisition.

That suspicion has a documented precedent in litigation involving Anthropic, where a federal court described millions of printed books being purchased, stripped of their bindings, scanned and discarded. The court treated those lawfully purchased books differently from a separate library of pirated copies, making acquisition method central to the copyright dispute. The case does not settle the legality of AI training broadly, but it shows why physical books can become useful when desired material is absent from convenient digital channels.Scale of Anthropic’s Early Book Acquisitions

Australian booksellers have reported similar patterns and a narrower preservation risk. Destructive scanning can save the words while eliminating physical evidence carried by a particular copy, including annotations or provenance. There is no evidence that AI companies systematically target rare books, but scarce material can lose historical context when the object itself is destroyed.

The unsettling image of a book surviving for decades only to lose its spine gives the economic shift a physical form. Yet the books are the warning, not the point. They reveal a new consumer assigning value to knowledge by different rules.

The Scale Behind AI Book Acquisition
Acquisition Signal Documented Scale
Anthropic print purchases Millions of books
Print acquisition spending Many millions of dollars
Pirated digital library More than 7 million copies
Reported UK seller orders About 6,000 books

Sources: United States District Court Northern District of California, The Guardian

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

When a Machine Values a Book Differently

Traditional secondhand markets price books around human interest, with readership and collectability often determining whether an old title retains demand. Artificial intelligence introduces another valuation question: does the book contain useful information that cannot be obtained easily elsewhere?

That question can reverse the economics of obscurity because low readership does not imply low informational value. A technical manual may preserve terminology that never migrated to the searchable web, while a local history may record facts barely represented in digital archives. Machine demand can therefore reward informational gaps that human markets largely ignored.

The unusual purchasing patterns become more coherent under that logic, even though no dataset yet establishes a market-wide AI premium for obscure books. A biography and an agricultural treatise need not share a subject if both fill missing areas in an existing corpus. Informational uniqueness can acquire value independently of the audience for the object containing it.

Older books may also carry provenance advantages because they predate the mass spread of generative AI across the information environment. That does not make older material inherently superior, and current evidence does not show a broad premium for pre-AI books. It does mean origin becomes more consequential as synthetic text grows easier to produce and harder to distinguish from human material.

Artificial intelligence can create text at immense scale while increasing the relative scarcity of differentiated information with credible provenance or specialist authority. By lowering the cost of finding and applying neglected knowledge, machines can make obscurity economically relevant rather than commercially fatal.

How Much Human Knowledge Remains Digitally Underrepresented
Measure Scale
Languages spoken worldwide 7,000+
Languages represented online About 1,000
Web content concentrated in 14 languages More than 91%
People still offline 2.6 billion

Sources: UNESCO, ICANN

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

The Return of the Information Silo

For much of the internet era, information trapped inside an institution or private database was treated as an inefficiency to overcome through digitization and search. Artificial intelligence changes part of that logic because useful knowledge can remain external to a model while still shaping its answers.

Retrieval-augmented generation allows an AI application to search an external collection and provide relevant material to a language model as grounding context. Microsoft describes the approach as a way to use proprietary or non-public information without embedding every useful document inside the model itself.

Economically, external retrieval separates ownership from use and creates room for controlled access rather than permanent transfer. An institution can sell a copy of a dataset, or it can keep the collection and allow machines to consult it under defined conditions. The second model preserves more control over the underlying knowledge.

Possession alone, however, does not create a valuable AI asset. An archive must become reliable enough for machines to use: scans need accurate text, documents need context, rights need clarity, and retrieval needs to work. UK guidance on AI-ready government data similarly treats ownership and provenance as foundations for machine use rather than administrative details.

A warehouse containing 20,000 anatomy books remains a library until the collection becomes searchable, trustworthy, legally usable and difficult to reproduce. At that point its economic character begins to change because value lies in controlled access to organized knowledge, not merely in possession of documents.

Digitized Does Not Mean Ready to Use
Europeana Measure Scale
Digital heritage items 59 million+
Data providers 3,500+
Content freely reusable 42%
Records raised to quality standard in one year 3 million

Sources: Europeana

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

The Next AI Advantage May Be What a Model Can Reach

Generative-AI competition has centered heavily on model performance and the computing resources required to produce it, but retrieval systems create another potential source of differentiation. Two models with similar general capabilities can produce different results when one can consult authoritative information unavailable to the other.

That asymmetry changes incentives for developers and information owners. Developers can acquire or license differentiated collections, while owners can reconsider whether transferring a permanent copy captures the full value of what they control. Firms may build defensible AI services around specialized knowledge without competing directly to build frontier models.

Reddit’s Long-Term Content Licensing Backlog
Reddit’s Long-Term Content Licensing Backlog

The relevant scarcity is often not the information itself because digital documents can be copied cheaply without depleting the original. Scarcity can instead reside in lawful machine-use rights, authenticated provenance or continuing access to a corpus that competitors cannot readily reproduce.

Copyright disputes reinforce the commercial importance of those distinctions even while broader training law remains unsettled. Anthropic’s litigation separated model training from the acquisition and retention of pirated books, making lawful provenance economically significant when information becomes an AI input.

A mature global market for privileged information access has not yet formed, and specialist knowledge does not replace model quality or computing power. As capable models spread, however, trusted information can become a stronger source of differentiation. Competitive advantage may increasingly depend on what a model is permitted to reach.

What Companies Are Paying for AI Data Access
Data Arrangement Reported Value Information Type
Google and Reddit About $60 million a year User discussions
OpenAI and News Corp Reported at $250+ million over 5 years Current and archived news
Shutterstock Big Tech deals Initially $25–50 million each Images, video and music
Defined.ai long-form video $100–300 per hour Licensed training media

Sources: Reuters

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Who Owns the Knowledge AI Has Not Yet Found

Institutions outside the frontier-model race may possess information poorly represented in current AI systems, including local knowledge or specialized research developed outside dominant digital channels. Material that sits at the edge of one information economy can become strategically useful in another.

Many institutions still treat such collections mainly as preservation responsibilities or administrative costs rather than as potential components of AI strategy. UK policy already separates possession from machine readiness by tying useful datasets to clear ownership, licensing and provenance.

The resulting policy problem resembles natural-resource economics only in a limited sense because information is not depleted when it is copied. Control can still shift when rights are transferred before owners understand future demand. Bargaining power may disappear even while the original documents remain physically or digitally intact.

Institutions therefore face choices about the terms of machine access rather than a simple choice between secrecy and unrestricted release. They can preserve originals while digitizing collections, distinguish public reading from commercial machine use, and license access without surrendering permanent possession.

For organizations sitting on decades of overlooked material, the emerging economic question is what becomes more valuable when machines begin searching systematically for what human markets neglected.

How Much Institutional Knowledge Is Still Outside Digital Reach
U.S. National Archives Measure Scale
Estimated textual pages 12.25 billion
Textual scans online 462 million
Estimated textual holdings digitized 3.78%
Electronic records held 33+ billion

Sources: U.S. National Archives and Records Administration

━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

The Books Are the Warning

The strange secondhand-book market gives that question a physical form because a low-value object can suddenly enter a different economic system. A book survives for decades, reaches a dealer with little human demand and then attracts a buyer interested primarily in whether its information exists elsewhere.

When the binding comes off and the pages become data, the object may disappear while its informational content acquires a new use. Most information silos will not announce that transition so visibly, since a scientific archive or corporate knowledge base can gain machine value without leaving its owner.

The broader shift is that information once difficult to find, organize or monetize can become more useful when machines reduce the cost of locating and applying it. The first AI race concentrated on systems capable of learning from enormous quantities of available information; a growing part of the next may concern access to what was never readily available.

Books disappearing into scanners are therefore more than a story about digitization or destruction. They show artificial intelligence beginning to alter the economic value attached to possessing human knowledge.

Measuring the AI Era Value of an Information Silo
Value Dimension Useful Metric Higher Value Signal
Information scarcity Share unavailable online Large offline share
Provenance Sources with verified origin High verification rate
Rights clarity Corpus with defined machine-use rights High rights coverage
Machine readiness Searchable, structured records High usable share
Replication difficulty Comparable external sources Few substitutes
Access control Permissioned retrieval share Granular control

Sources: UK Government Digital Service, Microsoft Azure


TL;DR Summary

  • Secondhand booksellers have reported unusual bulk purchases of obscure and apparently unrelated titles.
  • One British seller reported roughly 6,000 books ordered by buyers fitting the pattern since January 2026.
  • Anthropic’s litigation established a documented precedent for purchasing, destructively scanning and discarding printed books for AI development.
  • The evidence does not show that every unusual book order is connected to AI or that rare books are being systematically targeted.
  • AI can assign value to informational uniqueness even when conventional human demand for a book is weak.
  • Older books may carry additional provenance value because they predate the mass spread of generative AI content.
  • Retrieval-augmented generation allows models to use external proprietary information without permanently embedding every document in model parameters.
  • Controlled information silos can therefore become economically useful when their contents are reliable, searchable and legally usable.
  • Scarcity may increasingly lie in machine-use rights, provenance and controlled access rather than in the ability to copy digital information.
  • Firms controlling specialized knowledge may gain an AI advantage without competing directly to build frontier models.
  • Governments, libraries and universities may possess overlooked information whose future machine value exceeds its current market value.
  • Books disappearing into scanners provide an early physical signal that AI is changing the economics of possessing human knowledge.

Sources

  • The Guardian; Secondhand booksellers in UK and Ireland suspect AI firms behind strange bulk orders; – Link
  • The Guardian; More than just objects Australian booksellers raise alarm over destruction of rare titles to feed AI; – Link
  • U.S. District Court Northern District of California; Bartz et al v Anthropic PBC Order on Fair Use; – Link

When a Machine Values a Book Differently

  • British Library; 50 Facts About the British Library; – Link
  • British Library; Our Collections; – Link
  • Institute of Internet Economics; AI and the End of the Webpage; – Link

The Return of the Information Silo

  • Microsoft Azure; Retrieval Augmented Generation in Azure AI Search; – Link
  • UK Government Digital Service and Department for Science Innovation and Technology; Guidelines and Best Practices for Making Government Datasets Ready for AI; – Link
  • U.S. National Archives and Records Administration; Record Group Explorer; – Link

The Next AI Advantage May Be What a Model Can Reach

  • Reddit; Quarterly SEC Filing Revenue Contract Disclosures; – Link
  • Reuters; Inside Big Tech’s Underground Race to Buy AI Training Data; – Link
  • OpenAI and News Corp; A Landmark Multi Year Global Partnership with News Corp; – Link

Who Owns the Knowledge AI Has Not Yet Found

  • OECD; Enhancing Access to and Sharing of Data in the Age of Artificial Intelligence; – Link
  • OECD; Digital Government Index and Open Useful and Reusable Data Index 2025 Results and Key Findings; – Link
  • UK Department for Science Innovation and Technology; National Data Library Progress Update January 2026; – Link

The Books Are the Warning

  • UNESCO; Digitization Projects; – Link
  • U.S. Copyright Office; Copyright and Artificial Intelligence; – Link
  • U.S. District Court Northern District of California; Bartz et al v Anthropic PBC Final Approval of Class Action Settlement; – Link
  • Institute of Internet Economics; Authors Rally Against AI’s Unauthorized Use of Millions of Books; – Link

 

Keywords: Artificial Intelligence, Books, Knowledge, Information Silos, Machine Retrieval, Data Provenance, Knowledge Economics
Latest News

Can We Trust the Answer?

Can answers we find on the web be trusted? Can we trust that our favorite e-commerce site is showing...

More Articles Like This

- Advertisement -spot_img