Transcript
In early 2024, Anthropic was quietly acquiring millions of physical books, running them through hydraulic cutting machines, scanning the separated pages, and discarding the originals. The programme was called Project Panama. An internal document described it plainly: "Project Panama is our effort to destructively scan all the books in the world. We don't want it to be known that we are working on this." That document was unsealed in January 2026 as part of a copyright case against the company. This episode explains what the programme involved, how the law has treated it so far, and why the legal outcomes are more complicated than either side of the public debate tends to acknowledge.
The guiding question is this: was Project Panama a legally defensible acquisition of physical property, a legally reckless act of cultural destruction, or something in between? Based on the available evidence, it appears to be all three simultaneously — depending on which part of Anthropic's data strategy you examine. That tension is not rhetorical. It is built into the legal record itself.
Start with the timeline. According to the Washington Post's reporting based on those unsealed filings, Anthropic executives ramped up Project Panama in early 2024. The internal document's language — "we don't want it to be known" — suggests deliberate concealment, not simply routine confidentiality. The filings were unsealed in January 2026. Around six months later, in late July 2026, a post by the account @HedgieMarkets on X describing the programme reached approximately 21 million views, which is when the story crossed from tech-press coverage into mainstream public outrage. The underlying reporting had been done earlier that week by 404 Media.
The mechanics of the programme matter legally and practically. Anthropic — spending what Novara Media reports as tens of millions of dollars — acquired millions of physical books, primarily through secondary markets. A brokerage service called ISBNdb, according to the Dallas Express, now facilitates anonymous bulk orders of up to one million volumes specifically for AI clients. Pre-2022 titles commanded higher prices in this market, because books published before the widespread integration of AI-generated text are, by definition, free of synthetic content. That detail tells you something about what these companies were actually after: dense, human-authored prose at industrial scale.
Once the books arrived, they were processed through hydraulic cutting machines that removed spines and separated pages, which were then fed through high-speed industrial scanners. After digitisation, the physical originals were discarded — not archived, not donated. A secondary-market bookseller quoted by Novara Media described the arrangement as a useful way to clear hard-to-sell inventory, but also expressed concern about what happens when an uncommon title, with few surviving copies anywhere, enters this pipeline. That concern is not hypothetical. The sources confirm that rare editions have entered the process. What they cannot confirm — and this is one of the genuinely open questions — is which specific titles, in what quantities, and with what consequence for the broader archival record.
Now to the legal question, because this is where the story becomes structurally complicated rather than simply alarming.
A federal judge in the Northern District of California, U.S. District Judge William Alsup, ruled in June 2025, in a case called Bartz v. Anthropic, that digitising legally purchased physical books for AI training constitutes fair use. The reasoning was specific: Anthropic had bought the physical copies through legitimate channels, no new copies were redistributed, and the use was transformative. Under the first-sale doctrine — a long-established principle in U.S. copyright law that allows buyers to resell, lend, or otherwise use a purchased physical item without the original rightsholder's permission — owning the book means you can, in principle, do as you like with the object itself. The ruling said that scanning a legally purchased book and then training a model on the resulting text fell within that doctrine and within fair use.
That ruling is significant, and it is also genuinely contested. Supporters of the decision argue that it is consistent with how the first-sale doctrine has always worked: you bought the object, the object is yours. Critics point out that the economic effect on authors is the same whether the copying was technically legal or not — Anthropic extracted the commercial value of those books without compensating the people who wrote them. Both positions have weight.
Here is the complication the legal record makes impossible to ignore. At the same time that the physical scanning was ruled fair use, Anthropic separately agreed to a $1.5 billion class-action settlement — that figure comes from Novara Media's reporting — with authors whose work had been used to train Claude through a different channel: pirated digital copies, obtained without purchase or authorisation. The same training project, two different acquisition methods, two entirely different legal outcomes.
What this tells you is that the physical-book route appears to have been, at least in part, a legal strategy. By purchasing books outright and destroying them in the scanning process, Anthropic could argue — and a court accepted — that no unlicensed copying occurred. The book was bought. The book was scanned. The book no longer exists. Compare that to downloading pirated ebooks from shadow libraries, which left Anthropic holding digital copies it never had permission to make — copies that were straightforwardly infringing and that cost the company $1.5 billion to settle. The physical destruction was not incidental to the programme. It may have been the point.
Whether future courts will treat the combination of destructive physical scanning and separate use of pirated digital copies as a unified course of conduct — or continue to evaluate them independently — is unresolved. Some legal scholars will argue that copyright protection should consider cumulative harm to authors regardless of the technical method. Others will maintain that the physical and digital questions are legally distinct and must be assessed separately. The ruling in Bartz v. Anthropic does not settle this, and the settlement in the piracy case did not create a precedent that applies to the physical scanning question.
There is also the question of how this story spread and why it spread the way it did. The @HedgieMarkets post reaching 21 million views introduced the story to a large audience that encountered it framed around the most viscerally disturbing element: books being physically destroyed. That framing is accurate but incomplete. It tends to omit the fair-use ruling, the legal distinction between physical and digital acquisition, and the fact that the books being acquired were, in the main, commercially purchased secondhand copies rather than taken from libraries or private collections. The viral version of the story also merged concerns that are real — rare editions being permanently lost — with implications the evidence does not fully support, such as the suggestion that library collections were directly targeted. Nothing in the primary documents supports that specific claim. The concern about rare books is legitimate; the broader institutional library-raid framing appears to be an inference that outpaced the evidence.
That said, the concern about cultural loss is not nothing. A bookseller who spoke to 404 Media noted that uncommon books — titles with a small number of surviving physical copies — have entered this pipeline. The evidence is anecdotal on this point. There is no systematic accounting of what has been destroyed, and that is one of the most important gaps in the current record. The difference between a warehouse of surplus mass-market paperbacks and a run of regionally printed mid-century books with no digital equivalent is enormous, and the available sources do not draw that distinction clearly.
So here is where the record stands. Project Panama was a deliberately concealed programme to acquire and destroy physical books at scale for AI training. It ramped up in early 2024, became public through unsealed court filings in January 2026, and reached mass awareness in July 2026. A federal court ruled the physical scanning fair use under the first-sale doctrine. The same company paid $1.5 billion to settle claims over pirated digital works used for the same purpose. The legal split between those two outcomes reflects a genuine strategic choice: physical acquisition is legally cleaner, even if it is physically destructive. Whether other AI companies are running equivalent programmes is not confirmed in the current sources. What was destroyed, specifically, is not known. And how courts will eventually treat the combination of both acquisition strategies — physical and pirated digital — within a single training corpus remains an open question.
The mechanism at the centre of this is not mysterious once you see it clearly: AI companies need human-authored text at a scale that no licensing framework has been designed to accommodate, and Anthropic found a route through physical acquisition that a court would accept, while separately relying on pirated digital copies that a court would not. The legal system has so far treated those two routes as distinct. Whether they remain distinct, given that they fed the same model, is what the law has not yet resolved.