I love Anna’s archive
What if they turn around and encode all this data in crystals with advanced laser engraving aka making the information much more resilient to time and degradation? That said the sun will eventually run out of fuel and become a red giant and vaporize the Earth so putting a copy on a few other astrological bodies would be cool for the long haul.
You know what’s even cheaper than destrucive scanning of old books?
Scaring a bunch of volunteers into scanning them for free.
I’m not saying that was the plan all along. But it might have been. It’s plausible, with these jerks, isn’t it?
You do have to wonder who paid to get this content in front of everyone. How did anyone even know?
All part of the plan for these assholes to change history in their favor and teach AI to regurgitate their bullshit propaganda to the naive new generation. We are living literal 1984.
How to scan a book?
You take pictures of it with your phone. It is as simple as that.
https://archive.org/details/eliza-digitizing-book_202107
That’s how to do it the right way.
Damn that’s a lot of waiting. There was another device that just MRI’d books where they were, but no, ‘AI’ companies need to rip books apart to scan them. FU.
No less frustrating but I believe destruction is a required part of the process so the AI companies can say they still only have one “copy” of the book.
Oh like they care.
No, it’s far more crass: destroying the book makes scanning easier and faster.
Yah, exclusive content, or just rats watching the ship go down while burning books. FU AI bros playing the largest scam in history…
They have cradles that hold them so a camera can photograph them. You then save the book afterwards
The destruction of the book is done to mitigate some copyright concerns. There will be a fairly small number of large-scale projects like this, and each only needs one copy of a given book, so outside of maybe very rare books, the impact is going to be pretty negligible on availability — like someone buying one copy of that book.
Weird how when I buy a book, the “license” for the content is tied to the physical object, but when a company destroys the book and OCRs it, it’s perfectly ok
The destruction of the book is done to mitigate some copyright concerns.
Can you explain? The First Sale doctrine means its OK to sell books after you’re done with them. I would think it would be the scanning itself that could be a copyright problem, not the disposition of the physical book after.
I thought the real reason, the main reason, was, as the article says, “Destroying books is cheaper.”
I agree that the concern over the destruction of physical books, and what impact that will have on their availability, is probably exaggerated. Heck, the actual importance of old books to the AI companies may be exaggerated, that seems to be their hype style.
404 media reported this recently. Excerpt from the article :
It’s possible to scan books without destroying them, but cutting the spine makes it cheaper and faster. Additionally, the judge in the lawsuit ruled it was fair use and not a copyright violation for Anthropic to scan a book for training data in part because it destroyed the original, printed copy. Essentially, it’s the customer’s right to take physical media and store it digitally, and destroying the original copy means that copy isn’t duplicated and resold, and isn’t cutting into the publisher’s business.
Then they plagarize that data though, using it without attribution or payment, which is aganst copyright law. If a person did this they would be stopped. Because it’s big money, the law doesn’t apply.
Now you’ve entered the land of fairy tales and make believe. A person who memorizes a book and then burns it and then recites to others would in fact not be put in front of a firing squad like you want to.
Not saying corporations are people (they are though) but they are literally not doing anything else than that. Now they only have to teach the AI to learn double think, to learn from the book, but then forget the book, only to be able to remember what that book was about, while being unable to faithfully repeat it.
That’s literally going to solve any and all problems with AI. It’s not like we should demand meaningful regulations about the actual problems like authoritarian China does. You know, because, think of the book publishers. They are people too!
If a person reproduced copyrighted works and sold them, they could be sued under patent law. In the US, and in all of europe that I’m aware, at least in the west. In Japan, in South Korea, and a host of other countries.
Reciting a book is one thing, reproducing parts of it for commercial gain is another and is absolutely a violation of patent law. Just as a musician that takes a song and changes the lyrics and beat a bit and releases it would be sued.
Do you understand the difference?
Yeah, memorization is a bug they have to fix. Like I said, then all problems will be solved.
No problems will be solved in our favour. Us being those that aren’t the oligarchy and their well to do tools.
I’m surprised to hear someone on .ml defending llm’s. I thought you guys knew which side you were on on this issue?
These AA frauds are directly providing AI companies with scraped data, no?
No
Yes…
The document notes how Nvidia reached out to Anna’s Archive, stating that it’s “exploring including Anna’s Archive in pre-training data for [Nvidia’s] LLMs.
“Internal documents show competitive pressures drove Nvidia to piracy,” states the complaint, also revealing that before proceeding with the access, Anna’s Archive informed the company that its content was “illegally acquired and maintained.”
Despite this, Nvidia proceeded with the piracy, which resulted in the company receiving “millions of pirated copyrighted books,” or “roughly 500 terabytes of data.”
https://cybernews.com/ai-news/nvidia-annas-archive-ai-training/
That doesn’t sound like Annie Archive’s fault. Eveything has been scraped that they can access. If you run a free service, like wikipedia or AA, there is no way to stop them, as the law is in silicon valley’s pocket. No?
We can’t stop these ai companies at the point of piracy, not without the law. We have to find other ways, such as [redacted.]







