
Dr. Shivam Bhardwaj
assistant professor
The world's most modern machines are currently searching for the oldest source of knowledge. Companies that trained their models on the vast amount of public content available on the Internet now have access to millions of printed books. This change is not the result of a sudden respect for something old. This is a practical conclusion – the Internet may be full of information, but building better AI requires reliable knowledge more than information. This is where the story of Anthropic begins and this story is as interesting as it is uncomfortable.
Anthropic's case will not go down in history as the only copyright lawsuit or $1.5 billion settlement. It poses much bigger questions – why are books still essential despite the unlimited information available on the Internet, in whose hands will be the future of knowledge and how will public knowledge be protected in the age of artificial intelligence? It is possible that in the coming years, machines will start writing better than humans, but the foundation of how they will think is still being laid in those books, which have been written by humans, not machines. Perhaps the greatest paradox of the AI age is that the most intelligent machines of the future are becoming the disciples of the most patient writers of the past.
From Nalanda to Alexandria and Baghdad's Bait al-Hikma, history reminds us again and again that the destruction of books is not just the destruction of paper; This is an attack on the memory of civilization. Today we can scan a book and save every word in digital form and then destroy the original copy, but can the digital copy save the entire life of that book?
The smell of the book, its touch, the lines drawn by an old reader's pencil on the pages, a small note written on the margin and the feeling of keeping it in one's cupboard for years – none of these are fully captured in the data. A digital copy can preserve knowledge, but the original book also retains the human memory associated with that knowledge, so before giving books to AI, we need to ensure that the books survive for humans as well. Digitizing knowledge can be its preservation, but limiting it to private digital repositories can become not preservation but a new form of ownership.
After all, we don't want to leave for future generations a world where machines have our entire past and humans have only our data. It is not enough to save books to teach machines; They also have to save the humans to remember. The value of a serious book does not lie in its facts alone. Its real value lies in its intellectual structure—where to begin with a topic, which concepts are to be understood first, which arguments follow naturally from which arguments, and on what evidence the conclusions are drawn. A good book does not merely state its conclusion; She also often communicates opposing arguments and then establishes her position. This intellectual discipline is what makes it different from ordinary Internet content.
Another change in the last few years has made this challenge serious. A large portion of the new content published today is being created with the help of AI. In such a situation, if the next generation AI models will learn mainly from this content, then they will gradually start learning the patterns of the content written by machines. In the technical world, this is being seen not only as a crisis of copyright but also as a crisis of the quality of knowledge. It is in this context that fears of model lapse are discussed – that is, a situation when machines start losing their quality after repeatedly learning from machines. At such a time, literature written before the widespread dissemination of generative AI acquires special importance.
Documents revealed in a US court revealed that the company purchased a large number of printed books, scanned them and converted them into digital form. This project was internally named Project Panama. The scanning process itself was destructive – a hydraulic cutter would cut each book's cover, a high-speed scanner would digitize each page, and then the paper would be recycled. The physical book would be erased forever, leaving only its digital copy. That means a book was destroyed forever. The court considered this transformative digital conversion of purchased printed books to be fair use, as it was not a case of making new copies available in the market.
The same case also brought forth another truth, which received relatively less discussion. The hearing revealed that Anthropic had obtained more than seven million books from pirate websites even when there was the option to purchase them legally. The court deemed this illegal and the subsequent legal process approved a $1.5 billion settlement involving approximately 450,000 works of art. This made it clear that the court considered the scanning of purchased books and the use of pirated books as separate legal questions.
This is where this story becomes no longer just a copyright dispute. The real question isn't how Anthropic acquired the books; The real question is why did he need such a large number of books? At first glance the answer seems simple—books are more reliable, but the story goes deeper than that. If credibility were the only consideration, then research papers, government documents and certified websites would suffice. Then why would AI companies go after millions of books?