Artificial intelligence companies are competing for chips, computing power and billions of dollars in investment. But beneath that competition is another race that is less visible and potentially just as consequential: the race for knowledge.
That is what makes Project Panama so significant.
Project Panama was a confidential initiative developed by Anthropic, the company behind the AI chatbot Claude, to acquire large quantities of physical books and convert them into digital material for AI training.
The project was not simply about scanning books.
According to internal documents later unsealed in court, Anthropic described Project Panama as an effort to “destructively scan all the books in the world.” The plan involved acquiring physical books, removing or cutting their bindings, scanning their pages and recycling the physical copies.
The revelation offers an extraordinary glimpse into the lengths to which leading AI companies have gone in search of high-quality training data.
And it raises a question that extends far beyond copyright:
What happens when humanity’s accumulated knowledge becomes a strategic resource for private AI systems?
What Was Project Panama?

Project Panama emerged in early 2024 as Anthropic sought to expand the quantity and quality of material available to train its AI models.
The basic problem was straightforward. Large language models learn from enormous quantities of text. But the quality of that text matters. Anthropic wanted access to books because they provided long-form, structured and human-written material that could help its models develop stronger language capabilities.
According to documents reviewed by The Washington Post, Anthropic executives considered books particularly valuable for teaching Claude how to write well, rather than relying heavily on the increasingly repetitive and lower-quality language found across parts of the internet.
The solution was industrial in scale. Anthropic spent tens of millions of dollars acquiring potentially millions of books. The books were then prepared for high-volume scanning: their spines could be removed, their pages scanned and the physical copies discarded or recycled.
In other words, Anthropic was not primarily interested in preserving the books. It wanted the information inside them.
Why Destroy the Books?
At first glance, the destruction seems unnecessary. Why not simply scan a book and put it back on a shelf?
The answer is efficiency. Removing a book’s binding allows its pages to be processed rapidly through industrial scanning equipment. Instead of treating every book as a cultural object to preserve, the project treated it as a source of data to extract.
That distinction is important. For centuries, the physical book was the container of knowledge. In the AI economy, the physical object can become secondary.
Once its contents have been digitised, the pages can be transformed into data that can be stored, processed and used to train an AI model.
The book is the source. The data is the asset.
How Did Project Panama Become Public?
Project Panama was not announced by Anthropic.
Its existence emerged through a copyright lawsuit brought by authors against the company.
More than 4,000 pages of court documents eventually revealed details of Anthropic’s internal planning and its efforts to acquire and scan books on a massive scale. The documents included the internal description of Project Panama as an effort to destructively scan books and an instruction that Anthropic did not want the project to become known.
That secrecy became an important part of the controversy.
Anthropic was already facing legal challenges over the material it used to train its models. The Project Panama documents provided a much clearer picture of how aggressively the company had pursued books as a source of training data.
The Copyright Question Is Complicated
The controversy should not be reduced to the claim that Anthropic simply “destroyed copyrighted books.”
The legal picture is more complicated. In the underlying case, U.S. District Judge William Alsup ruled that Anthropic’s use of lawfully acquired books to train its AI models could qualify as fair use. The court viewed the training process as transformative rather than as a straightforward replacement for the original books.
But the case involved another, very different issue. Anthropic had also downloaded millions of books from unauthorized online “shadow libraries.” The court found problems with how those books had been acquired and stored. Rather than proceed to a full trial on those claims, Anthropic agreed to a $1.5 billion settlement with authors and other rights holders.
That distinction matters. The court did not simply rule that AI training on books is illegal.
Instead, the case highlighted a crucial difference between how material is acquired and how it is subsequently used for AI training.
The legal battle, however, is far from the end of the story.
Why Are Books So Valuable to AI?
The answer lies in a problem increasingly confronting AI developers: data quality.
The internet provided an enormous source of training material for early AI systems. But the online information environment is changing rapidly. Generative AI itself is producing vast amounts of text, creating concerns about repetitive or synthetic material entering future training datasets.
Books offer something different. They contain extended arguments, narratives, vocabulary, historical perspectives and forms of reasoning produced by human authors over decades and centuries.
For companies building increasingly capable AI systems, that makes books more than cultural products.
They are high-value training material. And that changes their economic significance.
From Books to Computational Capital
Project Panama reveals a striking transformation in how knowledge is valued.
A physical book can be purchased for a few dollars. But the information inside it can become part of a much larger computational system capable of generating text, answering questions, summarising research and producing new material.
Once digitised and incorporated into a model, the economic value of that information can extend far beyond the original book. This is why the debate should not stop at the destruction of physical copies.
The more consequential question is what happens after digitisation. The book disappears from the warehouse. The information does not.
The Emerging Power of Training Data
The AI race is usually framed around computing power.
But Project Panama points to another source of technological advantage: Who has access to the best information?
High-quality training data can influence what an AI model knows, how it performs and what it can produce.
That makes data increasingly valuable—not simply as a commercial asset, but as a source of technological power. The companies able to secure large and diverse collections of high-quality material may have an advantage over competitors that lack comparable access.
This is where the issue begins to move from technology into geopolitics.
The Geopolitics of Knowledge
The global AI competition is already reshaping relations between states.
Governments are competing over semiconductor supply chains, computing infrastructure, energy, talent and investment.
But there is another question emerging beneath these battles:
Who controls the knowledge that machines learn from?
A model trained predominantly on one linguistic, cultural or historical corpus may produce a different representation of the world from one trained on another. That matters.
AI systems are increasingly being used in education, journalism, research, business and government. Their outputs can influence how people understand events, societies and even one another.
The data that shapes those systems therefore has consequences beyond the technology sector.
It can influence whose knowledge is represented, whose history is visible and whose perspectives are reproduced at scale.
The Private Capture of Collective Knowledge
There is an uncomfortable paradox at the centre of Project Panama. The material being sought by AI companies represents centuries of collective human intellectual production. Literature. History. Science. Philosophy. Journalism. Culture.
Yet the infrastructure required to transform that knowledge into AI systems is concentrated among a relatively small number of private technology companies.
The knowledge may be collective. The models built from it are not.
That does not mean AI companies “own” human knowledge in any simple legal sense. Nor does Project Panama establish that Anthropic intended to control humanity’s intellectual record.
But it does reveal a structural shift: knowledge is increasingly being converted into computational infrastructure. And computational infrastructure creates power.
The Question After Project Panama
Project Panama is therefore about more than books.
It is about what happens when the accumulated intellectual output of humanity becomes an input into privately controlled AI systems.
The copyright debate will continue. Courts will decide what constitutes fair use. Authors and publishers will fight over compensation. AI companies will continue to seek new sources of training material.
But those arguments address only one part of the transformation.
The larger question is harder: Who decides what AI learns?, Which books are included?, Which languages?, Which histories?, Which voices?, Which archives?, And who has the resources to turn all of that knowledge into systems capable of shaping how millions of people access information?
Project Panama does not provide all the answers. But it offers a rare look at the scale of the competition.
The AI race is not only a race to build more powerful machines. It is also a race to acquire the knowledge that makes those machines more powerful. And in that race, the most important resource may not be the machine at all. It may be everything the machine has been taught.
Related stories:
Israel’s Million Dollar AI Campaign Against Gaza














