Anthropic Are Buying Rare Books, Feeding Them Into AI, Then Destroying Them

AI companies are quietly buying rare, human-written books, scanning them as training data, and destroying the originals in the process.

AI

Rare books and AI training data

Be the first to access our Large Action Model.

Get Paid to Train AI

Personal data is never stored

It sounds like the plot of a dystopian film.

But according to recently released court documents, internal company communications, and a reported $1.5 billion settlement, it's something that has already happened.

AI companies have quietly been collecting as many human-written books as possible to train their large language models, often without the consent of the authors or publishers. In the process, many of the original copies are being destroyed.

One internal Anthropic document outlined plans to "destructively scan all the books in the world."

Another point discussed keeping the project out of public view, stating:

"We don't want it to be known that we are working on this."

How on earth is this happening? Well, the answer is surprisingly simple. As AI companies race to build better models, they're all competing for the same thing: High-quality human-created data.

Quietly collecting the world's rarest books

AI companies have been collecting millions of human-written books to train their Large Language Models.

Many authors and publishers never gave permission for their work to be used this way. Some of these books were digitised using a process called 'destructive book scanning'.

A machine cuts the spine from the book, removes every page, and scans them individually at high speed.

The digital copy is kept. The original book is destroyed.

Why older books?

Anthropic reportedly focused on books published before 2022, before the rise of modern Large Language Models.

The reasoning was simple.

Books written before the AI era are guaranteed to contain entirely human-written content, making them especially valuable as training data.

As AI-generated content becomes more common, authentic human-created data is becoming one of the world's most valuable resources.


The lawsuit

Authors later filed a lawsuit against Anthropic, alleging the company had used millions of copyrighted books as AI training data without permission.

The court later issued a split ruling, distinguishing between different aspects of how the books were acquired and used.

Rather than continue the legal battle, Anthropic agreed to a reported $1.5 billion settlement, one of the largest copyright-related settlements associated with AI training.

The legal landscape is still evolving, but one thing is already clear:

AI companies need human-created data to build better AI.

But This story isn't really about books.

It's about ownership.

Books are just one example.

Every article.
Every image.
Every video.
Every line of code.
Every workflow.
Every action we perform online.

All of it has the potential to become training data for the next generation of AI.

Thirty years ago, oil was considered one of the world's most valuable resources.

Today, that resource is increasingly human-created data.

The question isn't whether AI should learn from humans. It already does.

The real question is:

If humans are creating one of the world's most valuable resources... shouldn't they have the opportunity to share in the value it creates?

Today, most AI models are built on human knowledge, creativity and actions.

Yet the people providing that value rarely own any part of the systems they're helping improve, or have much say over how their data is used.

We believe that should change.

That's why we built Action Model.

A community-owned AI alternative where people can contribute to training AI and participate in the value of the ecosystem they're helping build.

AI should not be owned by a few, it should be owned by all of us based on the value we helped create.

AI is inevitable.
Who benefits from it is still a choice.

Join our movement and combat Big Tech AI models:

https://train.actionmodel.com/sign-up

Table of content