Companies are purchasing physical books because books published before the widespread adoption of generative AI offer something...
...the contemporary web can no longer guarantee: text written by humans rather than by earlier AI systems. AI development is becoming ecologically, economically and intellectually unsustainable.
Summary: AI companies are buying and sometimes destroying pre-2022 books because they provide scarce, high-quality human-written training data uncontaminated by generative-AI output.
This risks removing rare works from circulation, enclosing cultural knowledge inside private datasets, bypassing creators and transferring publicly accessible material into opaque corporate systems.
As clean human data becomes scarcer, models may face higher costs, slower improvement and “model collapse” from repeatedly learning from synthetic output—making licensed, traceable and continuously updated human knowledge increasingly valuable.
The Last Clean Corpus: Why AI Companies Are Buying Old Books—and Why It Matters
by ChatGPT-5.6
The emerging market for old printed books as artificial-intelligence training material reveals a fundamental weakness in the present model-development paradigm. AI companies have built increasingly capable systems by absorbing enormous quantities of human-created text. Their own success is now contaminating one of their principal sources of future training data: the public internet.
The reported response is striking. Companies are purchasing physical books—sometimes in bulk and sometimes apparently for destructive scanning—because books published before the widespread adoption of generative AI offer something the contemporary web can no longer guarantee: text written by humans rather than by earlier AI systems.
This may appear to be an ingenious way of unlocking neglected knowledge. It is also a warning that the current approach to AI development is becoming ecologically, economically and intellectually unsustainable.
What is happening?
ISBNdb, traditionally a provider of book metadata to booksellers, libraries and distributors, is now marketing large-scale book-acquisition services to AI developers. Its central sales proposition is that printed books contain edited, structured and domain-specific human knowledge that cannot easily be replicated through ordinary web crawling. It specifically promotes books published before 2022 because they are assumed to be free from modern generative-AI content.
The company reportedly helps clients acquire between 1,000 and one million books per order, using ISBN information to locate titles, avoid duplication and organise the scanning process. It also promises strict confidentiality concerning the identities and acquisition strategies of its clients, while acknowledging the reputational problem created by headlines describing AI companies destroying millions of books.
There is a practical reason for destruction. High-volume scanning can be made faster and cheaper by cutting off a book’s spine and feeding the loose pages into scanning machinery. In the United States, a federal judge considering Anthropic’s book-digitisation programme treated the one-for-one replacement of a purchased physical copy with an internal digital copy as transformative fair use, emphasising that the original was destroyed and the digital version was not distributed externally. That fact-specific ruling should not be treated as a universal determination that every subsequent use of the digitised content is lawful, particularly outside the United States.
The article also reports unusual spikes in orders for rare, foreign-language and low-circulation books. One bookseller who normally sold around 20 books in a good week said that sales had risen to hundreds. However, the article appropriately acknowledges that the identities and purposes of the buyers were hidden and that the seller could not prove that every purchase came from an AI company. The evidence for a broader procurement trend is therefore suggestive rather than complete, although the existence of commercial book-sourcing services and Anthropic’s documented scanning programme establishes that the underlying practice is real.
Why the situation may be harmful
1. AI companies are destroying some of the cleanest surviving human data
The most obvious paradox is that the technology industry is physically consuming the very cultural material it regards as uniquely valuable.
A common book can usually be replaced. A low-circulation academic title, local history, minority-language work or out-of-print technical manual may have only a small number of surviving copies. The bookseller interviewed in the article warns that uncommon and foreign-language books may become materially harder to obtain if copies are destroyed during digitisation.
A digital scan is not a complete substitute for the original. It may omit bindings, marginalia, colour, illustrations, inserts, edition-specific features and evidence about the history of the object. It may also contain OCR errors or be stored in a proprietary dataset inaccessible to libraries, researchers and the public. Knowledge is thereby transferred from a publicly discoverable physical object into a privately controlled computational asset.
2. Purchasing a book is being treated as purchasing the knowledge economy around it
The argument that a second-hand book has already “discharged” its financial obligation to its creator is economically narrow. ISBNdb reportedly claims that secondary-market acquisition deprives authors of no income because the copies have previously been sold.
That reasoning overlooks the difference between reading one copy and using the text as an industrial input to produce a commercial model capable of serving millions of users. The model may substitute for books, reference works, professional advice or licensed databases. The economic value extracted from the work is therefore not necessarily exhausted by the original retail transaction.
Even where a particular form of copying is lawful, questions of remuneration, attribution, provenance and competition remain. A legally permissible activity can still produce an inequitable transfer of value from authors, publishers and cultural institutions to highly capitalised technology companies.
3. Secrecy prevents accountability
Confidential procurement arrangements make it difficult to know which books are being acquired, whether rare materials are being destroyed, how the resulting files are used and whether copyrighted works can later be removed or corrected.
Opacity also compromises dataset governance. A credible training-data system should record the identity and edition of each work, its source, applicable rights, permitted uses, scanning quality, corrections and eventual inclusion in particular models. A stack of books purchased through intermediaries and processed behind non-disclosure agreements may provide commercial discretion, but it does little to establish public trust.
The desire to avoid the “optics problem” is itself revealing. The concern appears to be less about whether destruction is culturally responsible than whether the public will see it happening.
4. The market can deprive libraries and readers of scarce works
AI companies have purchasing budgets that libraries, students, independent scholars and small cultural institutions cannot match. The bookseller quoted in the article observed both a disregard for price and a lack of thematic coherence in some bulk orders, which he interpreted as signs of machine-training acquisition rather than conventional collecting.
This can create a form of cultural enclosure. Materials that were dispersed throughout second-hand markets become aggregated into corporate datasets. The books may disappear from circulation, while the knowledge they contain remains accessible only indirectly through an AI product.
The result could be especially damaging for minority languages and poorly funded fields. The very works that can improve an AI model’s linguistic and cultural breadth may become less available to members of the communities that produced them.
5. “Pre-AI” does not necessarily mean accurate, representative or safe
Old books are attractive because they are assumed to be free of machine-generated text and deliberate anti-AI poisoning. The article describes them as “structurally clean” of modern contamination.
They are not epistemically clean. Historical collections include outdated science, discredited medical advice, racism, colonial assumptions, obsolete terminology, propaganda and factual errors. They may overrepresent societies and institutions that had the resources to publish books with ISBNs.
An ISBN-centred acquisition strategy introduces its own selection bias. It favours formally published, commercially catalogued works over oral traditions, unpublished archives, local publications, non-standard scripts and books predating widespread ISBN adoption. A model trained heavily on this material could become more literate while simultaneously becoming more historically anchored, culturally uneven and poorly informed about recent developments.
6. Digitisation introduces another layer of contamination
Printed text must be converted into machine-readable form. OCR errors can alter names, equations, tables, citations and non-Latin scripts. Page headers may be inserted into paragraphs; footnotes can become detached; illustrations may disappear; multiple columns can be read in the wrong order.
At million-book scale, apparently small error rates produce enormous quantities of corrupted text. Unless scanning is accompanied by stringent quality assurance, metadata and edition control, the supposedly pristine corpus may simply replace AI contamination with digitisation contamination.
7. The practice rewards the companies that polluted the web
There is a profound distributional unfairness in the emerging system. AI companies release tools that make automated content cheap and ubiquitous. The resulting output makes the open internet less reliable as a source of human-authored training material. The same companies can then use their capital to acquire the remaining protected reservoirs of human culture.
Human-created information becomes scarcer and more valuable precisely because machine-generated information has become abundant. The firms responsible for that shift are well positioned to capture the scarcity premium.
8. The physical and environmental costs are unnecessary
Bulk purchasing, international shipping, warehousing, destructive scanning, recycling, data storage and model training all consume energy and materials. Destroying books in order to create private digital copies is particularly difficult to defend where non-destructive scanning, publisher-supplied files or licensed digital archives could achieve the same purpose.
The process reflects a preference for acquiring control over data rather than building a sustainable knowledge infrastructure.
What this means for AI performance
The central technical concern is model collapse. Generative models learn a statistical representation of their training data. When their outputs are subsequently used to train new generations of models, small distortions can be reinforced. Rare patterns are lost, common patterns become disproportionately dominant and errors generated by one model are inherited by its successors.
Research published in Nature found that indiscriminate recursive training on model-generated data causes the less common “tails” of the original distribution to disappear. Over successive generations, models can lose diversity and increasingly misrepresent the underlying human data.
For language models, this can produce several performance problems:
Homogenisation. Models increasingly generate safe, predictable and statistically average language.
Loss of rare knowledge. Unusual events, minority perspectives, specialist vocabulary and uncommon linguistic constructions disappear first.
Error reinforcement. A plausible but false assertion generated by one model can appear in web pages, summaries and derivative documents, and later be treated as training evidence by another model.
Reduced originality. Recursive imitation encourages models to reproduce existing model conventions rather than the full variety of human expression.
Distorted confidence. Repetition across many AI-generated pages may make an unsupported claim appear well corroborated.
Weaker performance on edge cases. Models may retain competence on mainstream benchmarks while becoming worse at unusual, ambiguous or underrepresented tasks.
The risk should not be overstated. Synthetic data are not inherently harmful. Carefully designed synthetic examples can improve reasoning, coding, safety testing and performance in low-data domains. Collapse becomes most serious when synthetic material replaces human data, is repeatedly recycled without provenance, or is treated as independent evidence. Research indicates that retaining a meaningful supply of original data can substantially reduce degradation, while other work suggests that carefully controlled mixtures of real and synthetic data may avoid collapse.
What happens when the machines run short of human text?
The industry is unlikely literally to consume every available piece of text. The more plausible shortage concerns accessible, high-quality, legally usable, well-structured and reliably human-authored data.
One influential forecast estimated that, if historical scaling trends continued, the demand for public human-generated text could reach the available stock sometime between 2026 and 2032. This was a projection based on assumptions about model size and training practices, not a fixed deadline. Nevertheless, the purchasing of pre-2022 books suggests that some developers already perceive a scarcity problem.
As the easiest data sources are exhausted, several consequences are likely.
Model improvements will become more expensive
The cost of another trillion tokens will increasingly include licensing, digitisation, provenance checks, deduplication, human review and legal risk management. Compute may remain abundant for the largest companies, but trusted data will become the more restrictive resource.
Progress from brute-force scaling may slow
Models may continue improving through better architectures, inference-time computation, tool use and specialised training. But simply increasing the quantity of scraped text will offer diminishing returns. More data will not necessarily mean more knowledge when the new data largely consist of paraphrases, summaries and outputs from earlier models.
The frontier will become more concentrated
Companies with access to proprietary archives, search logs, workplace documents, scientific literature and large-scale human interaction data will enjoy a major advantage. Smaller laboratories and open-model developers may be left with increasingly contaminated public web corpora.
The result could be a two-tier AI ecosystem: highly capable models trained on licensed or privately collected human data, and weaker models trained primarily on recycled public material.
Current and newly created knowledge will become essential
Old books are useful for language, history and established domains, but they cannot describe scientific discoveries, laws, conflicts, cultural changes or technical developments that occurred after publication. Models require a continuing stream of new human observation, reporting, experimentation and interpretation.
An AI system cannot create reliable new facts merely by generating more sentences. New knowledge enters the world through research, measurement, journalism, scholarship and lived human experience. Undermining the institutions that produce these things ultimately undermines the future performance of AI itself.
A more sustainable direction
The appropriate response is not to prohibit the use of books in AI development. Books are valuable precisely because they embody sustained human thought, editorial investment and structured knowledge. They should be incorporated through practices that preserve both the works and the knowledge ecosystem that created them.
That would mean transparent and preferably non-destructive digitisation; direct licensing from authors, publishers and archives; work- and edition-level provenance; compensation for industrial reuse; preservation copies for libraries; robust OCR quality control; traceability from source to model; and mechanisms for correcting or withdrawing material.
Training strategies will also need to become more selective. Future progress will depend on smaller quantities of better data, deliberate mixtures of human and synthetic material, expert-created evaluation sets, retrieval from authoritative and updateable sources, and persistent access to the original human corpus.
The market for old books therefore tells us something larger than the price of training data. It shows that human-created knowledge is becoming the scarce natural resource of the AI economy. AI systems may be able to reproduce language at extraordinary scale, but they still depend on people and institutions to produce the reality that language describes. Destroying, enclosing or economically weakening that source may buy another generation of model improvement. It does not provide a sustainable foundation for the generations that follow.





