Anthropic, the American artificial intelligence company, has been revealed to have secretly destroyed hundreds of thousands of books to train its chatbot Claude. According to newly released court documents, the firm purchased large quantities of books, physically destroyed them, scanned their content, and used the data to improve its AI models. After digitization, the books were sent to recycling facilities where the paper was repurposed into cardboard packaging, paper products, and even toilet paper. The process described in the documents sounds almost like a scene from a dystopian novel. In massive warehouses, industrial machines cut through book spines, pages were separated one by one, and the contents were rapidly scanned. Within seconds, the material ended up in training databases for artificial intelligence. What remained of the books was nearly nothing, paper was sent to recycling plants where it was transformed into new materials. This operation, internally known as “Project Panama,” aimed to “destructively scan all books in the world.” One document states explicitly, “We don’t want the public to know we’re doing this.” The documents emerged during a lawsuit against Anthropic, initially reported by The Washington Post. They reveal a covert strategy to bypass traditional data sources and create more sophisticated AI models. Anthropic’s leadership concluded that the internet alone was insufficient for developing advanced AI. They wanted Claude to learn from professionally written books to develop a better writing style rather than relying on what they called “low-quality internet language.” Another reason cited was the growing issue of “AI model collapse”, a phenomenon where AI systems increasingly learn from content generated by other AIs, leading to a decline in quality over time. To achieve these goals, Anthropic sourced tens of thousands of books from multiple suppliers, including Better World Books and World of Books in the UK. External contractors handled the physical destruction and scanning. These contractors used machinery to remove book spines, separate pages, and quickly digitize the content. One partner noted that Anthropic planned to process between 500,000 and two million books within six months. The documents also expose the use of pirated books. The largest challenge for the company wasn't just destroying legally acquired books. Court filings show that co-founder Ben Mann had already downloaded copyrighted books via the pirate database LibGen in 2021. Evidence includes a video screenshot showing his computer screen during the download. A year later, he shared a link to Pirate Library Mirror, a vast collection of illegally distributed books, with a message: “Just in time!!!” In response to a collective lawsuit filed by authors who claimed Anthropic had used their works without permission to train its model, the company agreed to pay $1.5 billion in damages. This would cover approximately 500,000 copyrighted works. Attorney Justin Nelson of the authors' legal team stated this was the largest known copyright infringement settlement in history. However, the legal battle is far from over. Similar lawsuits are currently underway against Meta, OpenAI, Google, and Microsoft, alleging unauthorized use of protected works in AI development. A previous ruling by an American court suggested that using legally purchased books for AI training could qualify as permissible under U.S. law. However, this does not apply to millions of books obtained through illegal means, according to the plaintiffs. Anthropic maintains that the settlement did not change this part of the ruling and insists the dispute was solely about how certain materials were obtained, not the use of books themselves for AI training.
★
Keep the news honest.
ObjectiveNews is reader-funded and ad-free — we show you the bias instead of hiding it. Support independent journalism for €4/month.
Become a Supporter