Long live Anna’s Archive. I stand on the shoulders of the Internet, Wikipedia, Anna’s Archive, Z-Library, LibGen, YouTube, Hacker News, Reddit, and Sci-Hub.
I deeply admire the people who are obsessed with their passions and strive to build things that will lay the foundations for others.
Anthropic paying $1.5 billion in fines for downloading Anna's Archive established a moat. They want it to be illegal to pirate books: they can afford the penalties and continue doing it. Just like they want it to be illegal to run local ML inference.
OpenAI plainly admitted that it is impossible not to do so in a House of Lords inquiry. So, presumably there is no way around it to train models.
There is just not enough non-copyrighted data out there.
You mean this one https://committees.parliament.uk/writtenevidence/126981/pdf/ where they write "it would be impossible to train today’s leading AI models
without using copyrighted materials"? That doesn't mean they have to download those materials illegally. For a billion dollars, you can easily buy one legal copy of each book in Anna's Archive and still have some cash left over to run a whole-of-internet scraping operation.
I'm pretty sure we would know if they did that. And we don't.
Plus this is not legal in the EU (and Canada, and ... let's just say the entire rest of the world, and accept that I'll be wrong for one or two smaller countries). Doesn't that matter? Or is only Mistral disallowed from training on copyrighted materials? Je veux ma chaton fat, goddamit!
And where are you getting the idea that Mistral doesn't train on copyrighted data? There's not a lot of code written by people who've been dead for more than 70 years, but somehow Mistral has been able to release coding models anyway.
I wonder if anyone has run the numbers on what the actual cost, both in cash and logistical headache, contacting so many copyright holders would be. That seems like quite the feat to calculate.
There's an established network of intermediaries that can supply a large variety of books for a few dollars apiece, so no need to contact copyright holders directly.
This is very true. As someone with quite the experience with materials published under Penguin, Scholastic, etc. you effectively have a "dictionary attack" on the matter, rather than true "brute force," but that still leaves quite a list to compile to send to each and is easier for larger titles than smaller ones. I wonder how that leads to a bias in what materials get used for training. You are not getting many local self-published books this way.
It is almost like we need a "for use for training" agreement across the board. This would not fix the current issues (at least without substantial work), but going forward would allow for creators (or publishers/rights holders) such as this to designate a work as crawl-able for AI. A robots.txt just for Claude.
I thought the outcome of that was basically it's legal to train on books, but they acquired the books in the wrong way. If they went out and bought copies of them and trained it would have been fine
While it is definitely over a decade at this point (over two in fact), some of this likely comes from the term [citation needed], that originated on Wikipedia, as a cynical backhanded response to unsourced claims. It has become a catch-all. Language and how it evolves is a pretty interesting subject.
>Anthropic paying $1.5 billion in fines for downloading Anna's Archive established a moat. They want it to be illegal to pirate books: they can afford the penalties and continue doing it.
This seems like a "heads I win, tails you lose" type of argument. If Anthropic was pro-piracy I can imagine everyone getting mad that they're flouting law and want to "steal from artists" or whatever.
>and continue doing it
Source? AFAIK they were caught and stopped. That's why there was the recent story about how they were destroying old books to scan them.
> From the start, Anthropic “ha[d] many places from which” it could have purchased
books, but it preferred to steal them to avoid “legal/practice/business slog,” as cofounder and
chief executive officer Dario Amodei put it (see Opp. Exh. 27).
> they can afford the penalties and continue doing it.
I thought they could've bought just a single copy of each book and use the content to train their models. In that case, it falls into the fair use doctrine and they wouldn't need to pay the fine. And that will be way less expensive than the $1.5B price tag.
I'm pretty sure[0] they're all using shadow libraries, and saying things in favor of them would increase their liability.
Furthermore, every pirate wants to be an admiral. None of the big tech companies are actually in favor of any amount of copyright reform. They never have been. There is a huge gulf between "personally benefitting from copyright theft" and "actually wants to legalize the theft". Anthropic still believes they deserve to be paid for their models, they just have this delusion in their head that doing a bunch of computation on stolen data is equivalent to actual human creativity.
[0] OpenAI, Anthropic and Facebook have been shown in court to be using shadow libraries, I don't know about Google.
You missed Gigapedia (library.nu [1]) , which preceded most of the others.
Although I partially applaud what Anna's Archive is doing, I like their model the least (they want to CHARGE to download while doing Copyright infringement?).
In any case, I strongly believe that Anna's Archive is the wrong approach, as it has a single point of failure. We have been doing massive P2P sharing for more than 26 years; we have the algorithms for fully distributed file sharing and databases. Why are we still depending on HTTP/DNS based interfaces that are easily taken down by people wanting to limit knowledge?
I'm glad and thankful that the people behind Anna's Archive dedicate their time maintaining the huge base of human knowledge (Encyclopedia Galactica Asimov would say), but we (the people) should make it really distributed, really infallible and accessible (no, downloading 10TB torrent files doesn't make sense, except for archiving purposes).
We should have something like Popcorn Time but for knowledge.
> Although I partially applaud what Anna's Archive is doing, I like their model the least (they want to CHARGE to download while doing Copyright infringement?).
This is bollocks. AA gives users a means of paying to enjoy faster speeds as a means of contributing to costs, but the downloads are free to anyone who doesn't want to pay, and very often quick enough.
Whilst it pays for the service (and, in that respect, may be a necessary evil), it's morally questionable (at minimum) to charge for things that, by law, aren't yours in the first place.
Less morally questionable than claiming to users they are "buying" access to media that can be revoked at any point in the future with no recompense, of course referring to Sony and Amazon.
BBSes were the first, of course. In particular, Libgen, Sci-Hub and others can be traced back through several generations of libraries to the SU.BOOKS FidoNet echo conference created in the early 90's.
> Why are we still depending on HTTP/DNS based interfaces that are easily taken down by people wanting to limit knowledge?
Those are arguably the hardest protocols to block on the open Internet without causing major issues for all other sites, forcing those trying to take them down to play "whack-a-mole". If they were to create a new "AATP" for distributing data, it would make it trivial to block on every ISPs firewalls.
Probably the greatest achievement of modern society, a rebirth of the Library of Alexandria. Of course, only private for-business corporations are allowed to steal the world's knowledge, apparently, and only to be able to monetise it. There's something deeply wrong with our civilization that Anna's Archive is punished while companies like OpenAI, Google, Amazon, Anthropic, etc are just ... ignored when they do things like destructively digitise books or pirate things.
Exactly. All libraries are worthy of being supported and grown. I make no distinction between a physical dead-tree library and a digital library.
The only reason we even have dead-tree libraries at all is because 100 years ago, that was what John Rockefeller and Andrew Carnegie put forth to whitewash their horrible capitalist behaviors across the USA. And because it was done by those generations' billionaires, public libraries because acceptable.
If public libraries were created in the last 20 years, they would have been banned and felony copyright charges would have been levied. In fact, thats exactly what happened WHEN people tried to create free digital libraries.
> If public libraries were created in the last 20 years, they would have been banned and felony copyright charges would have been levied. In fact, thats exactly what happened WHEN people tried to create free digital libraries.
I wonder why nobody created a Netflix-like subscription for digital books
Amazon maybe? Kindle Unlimited. In some ways Audible with audiobooks. Which actually has lot of entrants though they tend to limit listening time on large part of their portfolio.
The problem with Libby is finding a local library that has a good selection. Mine does not, and my state does not have the feature of being able to get a library card from any library in the state. I would have to buy a membership to a non-local library, which can be surprisingly difficult to do if what you care about is a large selection of ebooks.
Because you have to pay a ton more for a lot of digital copies of books if you're doing any kind of lending and you're competing with existing libraries. So you both have higher costs (e.g. you have to pay more for a digital file that may have DRM, be required to self destruct after a certain number of uses, doesn't allow multiple lends at a time, etc.) and your prospective consumers have a free option so convincing them to pay is going to be difficult.
Netflix has a larger market base (far more people watch TV and movies than read books, particularly at a level high enough to justify a subscription) and a market base that is far less likely to already be aware of the free alternative (most serious readers of the sort that go through that amount of books are already aware of their local libraries, whereas serious movie/TV watchers are less likely to be so).
Additionally, modern Netflix is a streaming company, and even DVD Netflix was sending the materials directly to your house, which is a differentiator versus the library, whose materials need to be picked up. There's a convenience factor as an incentive to pay. Library streaming exists, but it's awful - very limited library and very limited watches - versus Netflix where once you sub you can watch as much as you want.
There was already one that existed in the web 2.0 era. Founded in 2012, raised a $3M seed from Founders Fund and a $14M A. Completely failed though and was acquihired by Google to have the founders lead Google Play Books.
Amazon now has Kindle Unlimited but I think it's mostly romance slop and self-published books. Seems like publishing rightsholders were just too inflexible to let the business model take off.
I consider libraries, public schools and maybe even public fire departments in the list of things that could never be proposed today if they didn't already exist.
I don't understand why there is so much love for Anna's archive here. I feel like I'm getting whiplash because there is so much hate for AI companies training and profiting off of the worlds knowledge without licensing it. And yet, a site that directly facilitates that by taking payments from AI companies is lauded as this amazing and honorable thing. Can someone help me understand what im missing?
In general I don't have many qualms with modern piracy, it just seems very hypocritical and I'm confused.
AI companies turn public data into closed commercial products. They are allowed to profit off your copyrighted data, but only they get to profit from their own models.
I'd imagine the backlash to AI companies from us white collar workers to be less severe if they have to publish their weights. In fact if you look closer you will see HN is actually pretty content with Chinese open models. It's the American AI corps with closed models that attract criticism.
I only wish they allowed browsing journals by year and volume. Libgen allowed that but war in Ukraine broken libgen. It lives but as a much shadier alter egos that does not support all that original libgen supported.
Even in the infamous Aaron Swartz case he was offered 6 months in prison as a plea bargain. Even ignoring the difference in crime (torrenting vs CFAA), no one was going to "lose most of your life".
when I read title .
"Annas archive owes $340 million absolutely first thing that came to mind is
" How much do all these Ai companies owe for their unauthorised use of books and other media?"
Anna's Archive is what TV told me in the 2000s that the future was. An online database with all published books one click away. Far from the dystopian reality than the corporate internet has become.
Just make them owe $134 Trillion or whatever the evaluation was years ago suggested by the RIAA for their estimation of 'damages' for music piracy. It's about as meaningful.
Might as well say "Annas Archive OWES ELEVENTY HUNDRED BILLIONTY-TRILLIONTY INFINITY DOLLARS!!!!!111" ala elementary school playground make-believe.
I remember when RIAA was shaking down 15 year olds for $7000 for a single Metallica download from Napster. Or how Aaron Schwartz was executed by proxy by JSTOR and the feds, for what should have been free to access for all.
But hey, Anthropic, OpenAI, X, and others can pirate to their hearts content with for-profit piracy, but "we" (royal) are OK with that. We just cant have the poors have access to the sum of human knowledge.
Many of us did, and I certainly still do, but there obviously weren't enough of us rebelling to convince them of how reprehensible it is to sue your own fans and sic the guns of the state on them.
Maybe if they use encryption, which would make them all pedophiles according to those governments. Hope you have your VPNs ready if visiting countries without a free internet, like North Korea, China, UK, or EU.
They provide the world with an invaluable service at the expense of their freedom. Seems incredibly noble and good to me. Especially considering the current climate.
Propaganda to what end? Are you referring to anything in particular?
Is your angle merely that shadow library hosts should be operating their services entirely for free (to the point of refusing payment) or is there something we're missing here?
It depends on what you consider to be a moral good. Some people feel that the world's information is getting locked down in the name of profit to the detriment of humankind. AA is providing knowledge to people who would otherwise not be able to get it. From that viewpoint, it's immoral to try to shut AA down.
OTOH, if you think that AA is robbing creators and publishers of their hard-earned proceeds, and that loss is greater than the loss to humanity from the destruction of AA, then you wouldn't think AA was a moral good.
I personally don't care about the morality, I'm just pointing out an absurd double standard.
You cannot claim that AI labs are bad for illegally downloading content to build their AI without paying creators, and that AA is good for illegally offering the same content.
I think both are immoral, but I definitely pirate all the content I can (except video games but only because it's unsafe, not to remunerate creators).
It is not grey at all, this is some bullshit that people tell themselves to feel better. Whether it's using the output of AI models or downloading a book on Anna's Archive, you are ultimately robbing the creator of profits.
Nobody on HackerNews would argue otherwise if their employer stole their code from their mind and didn't pay for it.
I agree that it's not grey at all, but because they're both so obviously good (particularly labs releasing open weight models). It's great that there are organizations making it ~free for anyone in the world to obtain information instantly. On the contrary, our policy of restricting something that naturally can be duplicated an unlimited number of times for free is obviously morally bad. People should be incentivized to create new ideas, not rent old ones for a century.
Of course on a related note, most work that's interesting to me was created by people who are dead anyway, so they don't mind.
Anyway, are you sure most people that support AA don't also support e.g. Deepseek and Qwen's efforts?
I think the reason why it shouldn't be claimed as a moral good is pretty obvious - authors deserve to get paid for their work. I certainly can't claim I stand on a bedrock of morality as I've used it from time to time, but even though I don't have the money that a lot of HN commenters do I make sure to not use it for small-time authors. Even really successful authors I'll only use it about 50% of the time.
If you want information to be free? Great. But most of those authors wouldn't be putting the work in to making that information/literature in the first place if they know they aren't going to get paid.
> But most of those authors wouldn't be putting the work in to making that information/literature in the first place if they know they aren't going to get paid.
My understanding is that the number of authors that can make a full-time job of writing is a rounding error compared to the population of authors. I believe that people should be paid for their work, but I don't think the current copyright regime is actually very effective at paying authors for their work. So I value copyright enforcement proportionately less based on that observation.
Meanwhile, on the flip side of the coin, copyright rulings are causing companies like Atheropic to destroy books as they scan them, creating a rising sense of panic around knowledge scarcity. This panic is a direct result of the scarcity of the copyright intended to create in the first place. It's entirely artificial.
I've been saying this for 25 years, but copyright is essentially broken.
Information itself should be free, authors deserve the right to monetise other ways (merch, physical media, exhibitions/shows, etc.) but society should pay artists and creators to do their thing[0]
Not everything needs to be a business, and I think art is one of those things
[0] I don't know much about it but maybe Ireland's Basic Income for the Arts scheme is a model. I think some other European countries do similar-ish things, too
Long live Anna’s Archive. I stand on the shoulders of the Internet, Wikipedia, Anna’s Archive, Z-Library, LibGen, YouTube, Hacker News, Reddit, and Sci-Hub.
I deeply admire the people who are obsessed with their passions and strive to build things that will lay the foundations for others.
Why don't Google, OpenAI, Anthropic, Facebook & Co defend Anna's Archive publicly?
Coming out would be a bold move for them.
Anthropic paying $1.5 billion in fines for downloading Anna's Archive established a moat. They want it to be illegal to pirate books: they can afford the penalties and continue doing it. Just like they want it to be illegal to run local ML inference.
In theory none of them actually got the right to train on illegally downloaded books. Anthropic was simply punished for doing it once.
One wonders if they're still doing it.
OpenAI plainly admitted that it is impossible not to do so in a House of Lords inquiry. So, presumably there is no way around it to train models. There is just not enough non-copyrighted data out there.
You mean this one https://committees.parliament.uk/writtenevidence/126981/pdf/ where they write "it would be impossible to train today’s leading AI models without using copyrighted materials"? That doesn't mean they have to download those materials illegally. For a billion dollars, you can easily buy one legal copy of each book in Anna's Archive and still have some cash left over to run a whole-of-internet scraping operation.
I'm pretty sure we would know if they did that. And we don't.
Plus this is not legal in the EU (and Canada, and ... let's just say the entire rest of the world, and accept that I'll be wrong for one or two smaller countries). Doesn't that matter? Or is only Mistral disallowed from training on copyrighted materials? Je veux ma chaton fat, goddamit!
AI companies legally acquiring books have indeed been in the news: https://news.ycombinator.com/item?id=49330742
And where are you getting the idea that Mistral doesn't train on copyrighted data? There's not a lot of code written by people who've been dead for more than 70 years, but somehow Mistral has been able to release coding models anyway.
But they have been training on copyrighted data since GPT-2 at least. 2019, and that's when it came out, so before that of course.
GPT-2 was trained using data scraped from the web (https://cdn.openai.com/better-language-models/language_model... section 2.1), i.e. copyrighted data provided free of charge to anyone with an internet connection.
I wonder if anyone has run the numbers on what the actual cost, both in cash and logistical headache, contacting so many copyright holders would be. That seems like quite the feat to calculate.
There's an established network of intermediaries that can supply a large variety of books for a few dollars apiece, so no need to contact copyright holders directly.
This is very true. As someone with quite the experience with materials published under Penguin, Scholastic, etc. you effectively have a "dictionary attack" on the matter, rather than true "brute force," but that still leaves quite a list to compile to send to each and is easier for larger titles than smaller ones. I wonder how that leads to a bias in what materials get used for training. You are not getting many local self-published books this way.
It is almost like we need a "for use for training" agreement across the board. This would not fix the current issues (at least without substantial work), but going forward would allow for creators (or publishers/rights holders) such as this to designate a work as crawl-able for AI. A robots.txt just for Claude.
I thought the outcome of that was basically it's legal to train on books, but they acquired the books in the wrong way. If they went out and bought copies of them and trained it would have been fine
> Just like they want it to be illegal to run local ML inference.
Citation?
https://news.ycombinator.com/item?id=49076057 ("Our position on open-weights models (anthropic.com)", 1812 comments)
Over the past decade I've noticed on HN the following order of frequency in choice of words, most common to least:
1. Citation
2. Source
3. Reference
Long ago in a career based on original research, I/we ONLY used "reference."
While it is definitely over a decade at this point (over two in fact), some of this likely comes from the term [citation needed], that originated on Wikipedia, as a cynical backhanded response to unsourced claims. It has become a catch-all. Language and how it evolves is a pretty interesting subject.
You're right. Has to be of Wikipedia use origin. Thanks!
>Anthropic paying $1.5 billion in fines for downloading Anna's Archive established a moat. They want it to be illegal to pirate books: they can afford the penalties and continue doing it.
This seems like a "heads I win, tails you lose" type of argument. If Anthropic was pro-piracy I can imagine everyone getting mad that they're flouting law and want to "steal from artists" or whatever.
>and continue doing it
Source? AFAIK they were caught and stopped. That's why there was the recent story about how they were destroying old books to scan them.
> From the start, Anthropic “ha[d] many places from which” it could have purchased books, but it preferred to steal them to avoid “legal/practice/business slog,” as cofounder and chief executive officer Dario Amodei put it (see Opp. Exh. 27).
https://cdn.arstechnica.net/wp-content/uploads/2025/06/Bartz...
Sounds pretty pro piracy to me.
> they can afford the penalties and continue doing it.
I thought they could've bought just a single copy of each book and use the content to train their models. In that case, it falls into the fair use doctrine and they wouldn't need to pay the fine. And that will be way less expensive than the $1.5B price tag.
That's what they're doing now, when they are established.
But when it was a proof of concept, they were using pirated data.
Just like Spotify did.
I'm pretty sure[0] they're all using shadow libraries, and saying things in favor of them would increase their liability.
Furthermore, every pirate wants to be an admiral. None of the big tech companies are actually in favor of any amount of copyright reform. They never have been. There is a huge gulf between "personally benefitting from copyright theft" and "actually wants to legalize the theft". Anthropic still believes they deserve to be paid for their models, they just have this delusion in their head that doing a bunch of computation on stolen data is equivalent to actual human creativity.
[0] OpenAI, Anthropic and Facebook have been shown in court to be using shadow libraries, I don't know about Google.
Those are all law-abiding organizations, which AA is not.
INB4: "Here is one time one of those organizations broke the law". Don't go there, absolute lowest level of conversation.
You missed Gigapedia (library.nu [1]) , which preceded most of the others.
Although I partially applaud what Anna's Archive is doing, I like their model the least (they want to CHARGE to download while doing Copyright infringement?).
In any case, I strongly believe that Anna's Archive is the wrong approach, as it has a single point of failure. We have been doing massive P2P sharing for more than 26 years; we have the algorithms for fully distributed file sharing and databases. Why are we still depending on HTTP/DNS based interfaces that are easily taken down by people wanting to limit knowledge?
I'm glad and thankful that the people behind Anna's Archive dedicate their time maintaining the huge base of human knowledge (Encyclopedia Galactica Asimov would say), but we (the people) should make it really distributed, really infallible and accessible (no, downloading 10TB torrent files doesn't make sense, except for archiving purposes).
We should have something like Popcorn Time but for knowledge.
[1] https://en.wikipedia.org/wiki/Library.nu
> Although I partially applaud what Anna's Archive is doing, I like their model the least (they want to CHARGE to download while doing Copyright infringement?).
This is bollocks. AA gives users a means of paying to enjoy faster speeds as a means of contributing to costs, but the downloads are free to anyone who doesn't want to pay, and very often quick enough.
Whilst it pays for the service (and, in that respect, may be a necessary evil), it's morally questionable (at minimum) to charge for things that, by law, aren't yours in the first place.
Less morally questionable than claiming to users they are "buying" access to media that can be revoked at any point in the future with no recompense, of course referring to Sony and Amazon.
BBSes were the first, of course. In particular, Libgen, Sci-Hub and others can be traced back through several generations of libraries to the SU.BOOKS FidoNet echo conference created in the early 90's.
> Why are we still depending on HTTP/DNS based interfaces that are easily taken down by people wanting to limit knowledge?
Those are arguably the hardest protocols to block on the open Internet without causing major issues for all other sites, forcing those trying to take them down to play "whack-a-mole". If they were to create a new "AATP" for distributing data, it would make it trivial to block on every ISPs firewalls.
Incredible website. Really one of the dreams of the internet realised, the ability to access all knowledge at your fingertips.
I was blown away when I paid to get an API key for it. The process is intricate and state of the art.
Probably the greatest achievement of modern society, a rebirth of the Library of Alexandria. Of course, only private for-business corporations are allowed to steal the world's knowledge, apparently, and only to be able to monetise it. There's something deeply wrong with our civilization that Anna's Archive is punished while companies like OpenAI, Google, Amazon, Anthropic, etc are just ... ignored when they do things like destructively digitise books or pirate things.
Didn’t they have a default judgment against them because they didn’t show up to court? Do they even know who runs it?
They railroaded Aaron Swartz (which eventually lead to his suicide) for much, much less.
One side wants to freely share knowledge with all of humanity, the other wants to restrict it to make a buck. I know who I support.
Anna's Archive is a gift to humanity
Exactly. All libraries are worthy of being supported and grown. I make no distinction between a physical dead-tree library and a digital library.
The only reason we even have dead-tree libraries at all is because 100 years ago, that was what John Rockefeller and Andrew Carnegie put forth to whitewash their horrible capitalist behaviors across the USA. And because it was done by those generations' billionaires, public libraries because acceptable.
If public libraries were created in the last 20 years, they would have been banned and felony copyright charges would have been levied. In fact, thats exactly what happened WHEN people tried to create free digital libraries.
> If public libraries were created in the last 20 years, they would have been banned and felony copyright charges would have been levied. In fact, thats exactly what happened WHEN people tried to create free digital libraries.
I wonder why nobody created a Netflix-like subscription for digital books
Amazon maybe? Kindle Unlimited. In some ways Audible with audiobooks. Which actually has lot of entrants though they tend to limit listening time on large part of their portfolio.
Libby very much exists, you just need to sign up through your local library.
The problem with Libby is finding a local library that has a good selection. Mine does not, and my state does not have the feature of being able to get a library card from any library in the state. I would have to buy a membership to a non-local library, which can be surprisingly difficult to do if what you care about is a large selection of ebooks.
Because you have to pay a ton more for a lot of digital copies of books if you're doing any kind of lending and you're competing with existing libraries. So you both have higher costs (e.g. you have to pay more for a digital file that may have DRM, be required to self destruct after a certain number of uses, doesn't allow multiple lends at a time, etc.) and your prospective consumers have a free option so convincing them to pay is going to be difficult.
The economics aren't there.
Public libraries lend DVDs and Blu-Rays yet Netflix still exists.
Netflix has a larger market base (far more people watch TV and movies than read books, particularly at a level high enough to justify a subscription) and a market base that is far less likely to already be aware of the free alternative (most serious readers of the sort that go through that amount of books are already aware of their local libraries, whereas serious movie/TV watchers are less likely to be so).
Additionally, modern Netflix is a streaming company, and even DVD Netflix was sending the materials directly to your house, which is a differentiator versus the library, whose materials need to be picked up. There's a convenience factor as an incentive to pay. Library streaming exists, but it's awful - very limited library and very limited watches - versus Netflix where once you sub you can watch as much as you want.
https://en.wikipedia.org/wiki/Oyster_(company)
There was already one that existed in the web 2.0 era. Founded in 2012, raised a $3M seed from Founders Fund and a $14M A. Completely failed though and was acquihired by Google to have the founders lead Google Play Books.
Amazon now has Kindle Unlimited but I think it's mostly romance slop and self-published books. Seems like publishing rightsholders were just too inflexible to let the business model take off.
I consider libraries, public schools and maybe even public fire departments in the list of things that could never be proposed today if they didn't already exist.
Public libraries pay for their books
Exactly. if AI gets a pass on stealing all of the world's information and gets a pass why shouldn't we get to enjoy the same benefits?
And Google owes 20 decillion dollars to the Russian government. Who cares either way.
I don't understand why there is so much love for Anna's archive here. I feel like I'm getting whiplash because there is so much hate for AI companies training and profiting off of the worlds knowledge without licensing it. And yet, a site that directly facilitates that by taking payments from AI companies is lauded as this amazing and honorable thing. Can someone help me understand what im missing?
In general I don't have many qualms with modern piracy, it just seems very hypocritical and I'm confused.
AI companies turn public data into closed commercial products. They are allowed to profit off your copyrighted data, but only they get to profit from their own models.
I'd imagine the backlash to AI companies from us white collar workers to be less severe if they have to publish their weights. In fact if you look closer you will see HN is actually pretty content with Chinese open models. It's the American AI corps with closed models that attract criticism.
Rembember to seed torrent kids
I wonder what would happen if Annas archive announced
In response to the 340M fine, .... "We have set up a LLM and are deleting all our books ? "
I only wish they allowed browsing journals by year and volume. Libgen allowed that but war in Ukraine broken libgen. It lives but as a much shadier alter egos that does not support all that original libgen supported.
So when are Nvidia, Facebook, and everyone else going to chip in and pay up?
Did you miss that they are huge corporations that the law doesn't apply to?
if this is true. what does this say about "rule of law"? is it all fiction?
You mean how anthropic lost a court case and now they're being made to pay?
they are paying from the revenue, they gained ton of valuation and revenue because of books? if you do that, you lose most of your life
Even in the infamous Aaron Swartz case he was offered 6 months in prison as a plea bargain. Even ignoring the difference in crime (torrenting vs CFAA), no one was going to "lose most of your life".
I don’t get it. I am not a corp. i can download books from AA (who does the law apply to exactly?)
when I read title . "Annas archive owes $340 million absolutely first thing that came to mind is " How much do all these Ai companies owe for their unauthorised use of books and other media?"
"owes" seems like a wrong word here
What are current domains? I tried the ones from article but got redirected to spam
.gd works for me
Anna's Archive is what TV told me in the 2000s that the future was. An online database with all published books one click away. Far from the dystopian reality than the corporate internet has become.
"If you own someone $340, that's your problem. If you owe someone $340,000,000, that's there problem."
Might as well be a trillion dollars as it's never getting paid.
Unfortunately it’s blocked in Germany
only the local providers in Germany are forced to block it via DNS - 8.8.8.8 to the rescue
Tor Browser is handy, if simply changing your DNS isn't enough
Just make them owe $134 Trillion or whatever the evaluation was years ago suggested by the RIAA for their estimation of 'damages' for music piracy. It's about as meaningful.
taking shots against the King?
better not miss!
Thanks for reminding me I need to download some books for the holidays!
"owes"
Might as well say "Annas Archive OWES ELEVENTY HUNDRED BILLIONTY-TRILLIONTY INFINITY DOLLARS!!!!!111" ala elementary school playground make-believe.
I remember when RIAA was shaking down 15 year olds for $7000 for a single Metallica download from Napster. Or how Aaron Schwartz was executed by proxy by JSTOR and the feds, for what should have been free to access for all.
But hey, Anthropic, OpenAI, X, and others can pirate to their hearts content with for-profit piracy, but "we" (royal) are OK with that. We just cant have the poors have access to the sum of human knowledge.
> I remember when RIAA was shaking down 15 year olds for $7000 for a single Metallica download from Napster.
Why didn't Metallica's fans rebel? That would have stopped it quickly and set a precedent and example for the rest.
Because they are a bunch of mask wearing vaccine lovers obviously
Many of us did, and I certainly still do, but there obviously weren't enough of us rebelling to convince them of how reprehensible it is to sue your own fans and sic the guns of the state on them.
Reminder for me to make another donation.
Piracy is morally justified at this point.
I wonder by if the US considers this “supporting a terrorist organization”.
Donate a significant amount of money, and you can get a free trip to the new gulag in El Salvador. Or maybe gitmo.
US probably considers anything impairing profits, stock prices, or enterprise value terrorist activity at this point.
Maybe.
Thankfully I don't live in the US, and likely won't set foot there ever again.
I wonder if the EU or UK does.
Maybe if they use encryption, which would make them all pedophiles according to those governments. Hope you have your VPNs ready if visiting countries without a free internet, like North Korea, China, UK, or EU.
Yeah, truthfully. This type of article is generally a nice nudge to donate.
[1] https://stallman.org/articles/end-war-on-sharing.html
I really don't understand why HackerNews gets hard for Anna's Archive but you never see other illegal download websites praised here.
They're not more or less virtuous than any others, they're in it for the money and you're a fool if you think otherwise.
I'm not against piracy, it's great, but to claim they are moral saints is a joke.
Anyway, bring in the downvotes as I know will happen.
As a hacker News viewer I'm for other illegal download sites as well.
If you just Google 'free media heck yeah' you'll see a great deal of them!
Yeah me too, but I don't claim that they're a moral good.
They provide the world with an invaluable service at the expense of their freedom. Seems incredibly noble and good to me. Especially considering the current climate.
They're a for profit organization and you're falling for their propaganda.
Propaganda to what end? Are you referring to anything in particular?
Is your angle merely that shadow library hosts should be operating their services entirely for free (to the point of refusing payment) or is there something we're missing here?
No, just that they shouldn't claim that they have a vertous mission and that this is what your money is going towards.
Freeing all information is virtuous, though…
It depends on what you consider to be a moral good. Some people feel that the world's information is getting locked down in the name of profit to the detriment of humankind. AA is providing knowledge to people who would otherwise not be able to get it. From that viewpoint, it's immoral to try to shut AA down.
OTOH, if you think that AA is robbing creators and publishers of their hard-earned proceeds, and that loss is greater than the loss to humanity from the destruction of AA, then you wouldn't think AA was a moral good.
Of course, it's all gray in the middle.
I personally don't care about the morality, I'm just pointing out an absurd double standard.
You cannot claim that AI labs are bad for illegally downloading content to build their AI without paying creators, and that AA is good for illegally offering the same content.
I think both are immoral, but I definitely pirate all the content I can (except video games but only because it's unsafe, not to remunerate creators).
It is not grey at all, this is some bullshit that people tell themselves to feel better. Whether it's using the output of AI models or downloading a book on Anna's Archive, you are ultimately robbing the creator of profits.
Nobody on HackerNews would argue otherwise if their employer stole their code from their mind and didn't pay for it.
I agree that it's not grey at all, but because they're both so obviously good (particularly labs releasing open weight models). It's great that there are organizations making it ~free for anyone in the world to obtain information instantly. On the contrary, our policy of restricting something that naturally can be duplicated an unlimited number of times for free is obviously morally bad. People should be incentivized to create new ideas, not rent old ones for a century.
Of course on a related note, most work that's interesting to me was created by people who are dead anyway, so they don't mind.
Anyway, are you sure most people that support AA don't also support e.g. Deepseek and Qwen's efforts?
We don't have to have the same exact moral systems of course, but this is curious: if we (collectively!) support the thing, why not claim it's good?
I think the reason why it shouldn't be claimed as a moral good is pretty obvious - authors deserve to get paid for their work. I certainly can't claim I stand on a bedrock of morality as I've used it from time to time, but even though I don't have the money that a lot of HN commenters do I make sure to not use it for small-time authors. Even really successful authors I'll only use it about 50% of the time.
If you want information to be free? Great. But most of those authors wouldn't be putting the work in to making that information/literature in the first place if they know they aren't going to get paid.
> But most of those authors wouldn't be putting the work in to making that information/literature in the first place if they know they aren't going to get paid.
My understanding is that the number of authors that can make a full-time job of writing is a rounding error compared to the population of authors. I believe that people should be paid for their work, but I don't think the current copyright regime is actually very effective at paying authors for their work. So I value copyright enforcement proportionately less based on that observation.
Meanwhile, on the flip side of the coin, copyright rulings are causing companies like Atheropic to destroy books as they scan them, creating a rising sense of panic around knowledge scarcity. This panic is a direct result of the scarcity of the copyright intended to create in the first place. It's entirely artificial.
I've been saying this for 25 years, but copyright is essentially broken.
> authors deserve to get paid for their work
I don't really agree
Information itself should be free, authors deserve the right to monetise other ways (merch, physical media, exhibitions/shows, etc.) but society should pay artists and creators to do their thing[0]
Not everything needs to be a business, and I think art is one of those things
[0] I don't know much about it but maybe Ireland's Basic Income for the Arts scheme is a model. I think some other European countries do similar-ish things, too
Can't talk for all of HN, but I for one hereby praise most/all piracy websites. Anna's Archive is great, and Rutracker is also great.
It's not a competition, and being a moral saint isn't a good KPI; being a net positive for society is.
I praise all pirates and leakers
Maybe you can own (some) things, you can't own information
it "Owes" just like Anthropic and OpenAI "Own" their models trained on the worlds collective IP.