Skip to main content
The 'AI' Problem Noone Is Talking About

The 'AI' Problem Noone Is Talking About

#books

#llm

#rights

#ai

We are very deeply into ‘AI’ and the environment it has created. Over the course of these several years there have been a number of discussions, debates on ethics and the proliferation of these technologies. It started with speak of them taking jobs - people getting fired, companies replacing humans with them. And more importantly the creativity side and generative ‘AI’.

To add to that, the legal debates on rights and who own what is created by these generative systems.

More recently, the effect this technology has had on the open-source world.

However, throughout all of this what I’ve never seen being talked about most is the core that powers these LLMs, what enables them to be functional and to exist, and that is the data.

For your ChatGPT, your Claude, your Gemini; for your Deepseek. For all of them to function and give you answer as ‘quickly’ as you assume, it requires data - large swathes of data. But there is only so much data one can feed into an LLM to train it. As of a certain point these companies had scrapped most if not all of the internet to train their model - some of it legally and some illegally. It had gotten to a point where the ‘AI’ was being trained on data from the ‘Ai’ creating a recursive loop.

So where does a multi-million dollar company turn to in order to satisfy the endless hunger of the beasts they created? Well, to books of course. Physical books.

A good number of you would not know about this because books, in the current climate, are not something that makes headlines. Also, a lot of you don’t care about books anymore. And this is why I have written this - the books.


How Are Books Relevant?

Remember I mentioned legal battles over ‘AI’? That’s where books come in. If these companies were truly altruistic, the would have gone the legal route, gotten the rights, paid whichever party had ownership of the material they wanted to use. Buuut, as has been evidence - and the reason for the legal battles - they have been scanning and stealing content they have not licensed.

For the books they have now turned to, they would also have to pay royalties and licensing fees…

In the backdrop of legal battles over ownership and who owns what is created by ‘AI’, books were being destroyed.

I doubt any of you would be aware of this as unless you explore, the only ‘AI’ you know is ChatGPT. But let me bring your attention to Anthropic and what its been doing…

Anthropic’s Destruction of Books

So as to not pay the fees and interact with the authors, Anthropic went about procuring old books, scanning them in a destructive meaner. In an effort to improve its LLM, in an initialve they called “Project Panama”

And I quote:

“Project Panama is our effort to destructively scan all the books in the world.”

Why do this you may ask? I have already mentioned why, but let me elaborate further.

Anthropic could have secured copyright permission to use existing e-books. This would have required going through the proper channels and engaging with the authors - something the co-founder described as ‘tedious’. They had already pirated some sources already which resulted in a large settlement. So, when going about it legally became ‘complicated’, they turned to destructively scanning books.

You can read more about the whole issue here

Of course they would rather destroy printed copies than work with their authors

The Loss Of Historical Data

A lot of the books being scanned and destroyed are one of a kind, can’t be replicated. Once they are gone, they are gone. The only remaining record would be locked in Anthropic’s servers. They have no obligation share what they have scanned, and most likely won’t. That will make it even more difficult to find and get actual sources from the original authors of said information. All that would be left is whatever output the LLM has produced from its training


What Happens When There Are No More Books?

We have already seen examples of what that looks like. Every single piece of physical media that was made digital exists at the whims of the company that made it so, and when that licensing runs out. In particular, the PlayStation situation - where hundreds of movies ‘purchased’ on the platform are slated for deletion.

While a lot of you might not care now, you will later on when you can’t get actual sources. When all you have to look for is an LLM run by a company that can revoke your access at their whim, at their own discretion. You can’t say anything about it because they’re in control .

In a world where all the physical books are gone and one company or another controls your access to those digital copies, nothing stops them from modifying the books, taking the books from your control, or just denying you access.

That is not a concern if you have physical copy of the book.


Wrapping Up

Do I have a solution? I don’t. I am not in a position to solve these things, not in a high enough position of power. All I can do is do what I’ve always been doing: try to bring awareness to what’s happening while you are asking ChatGPT what the square root of a triangle is.

All of this conveniently ties into the current push for digital ID that has been ramping up as of late. Once that goes online and the media is also online, the ID could be used to stop you from reading what they don’t want you to. I say this in full knowledge that a good number of you don’t read anyway, but for the ones that still do…

The domino effect concerns me a great deal…


Thank you for reading, let’s connect!

Thank you for visiting this little corner of mine. My email and DMs are always open; if you want to chat or collaborate on something, shoot me a email

Where you can get hold of me:

Back to articles