A one time payment like 1.5B doesn’t do anything. There needs to be a royalty payment based on if the AI regurgitates existing ideas. That is probably the correct way to legislate this. If anything a human does can instantly be copied by an LLM, and then sent to all its subscribers, things need to change
Yeah, it's always interesting the two-sides of a situation like this. Add regulation/enforcement to the big companies and you often shut out the smaller ones following.
Meta also has copyright lawsuits for the open models they released, so open models are not immune.
... unless the line we want to draw is "american orgs pay, others don't", as currently seems to be happening.
> Add regulation/enforcement to the big companies and you often shut out the smaller ones following.
That is the case, any regulation increases the cost to enter a market.
But in this case, its irrelevant because the moat of cost to enter is already unfathomable and secondly, they are not adding regulation but fining them for committing a crime.
So yeah, adding that every food compnay needs 3 health inspectors that they pay for would benefit coca cola over you mom and pop bakery. But telling someone they cannot start a Space agency with money laundered from ransom and drug sales payments would not affect much the competition markets
There are multiple ways to respond to that and I will try and summarise them.
Current believe is that its a "winner takes all market", so companies are acting rationally and using Brute Force compute to get there first. Training costs scale linearly, which means the moat is directly related to compute cost
There are theories that they are wasting 90% of training costs and there are more efficient ways to do it than throw compute at the problem. But if thats the case then chances are the market is not "winner takes all". Which then means the valuation of the ENTIRE market is overvalued.
Basically the only way for the assertion "at the moment" to be true is if the market is a bubble, else if the current theory of winner takes all market means a monopoly will make it so that cost isnt even the worst of the moats to enter.
They're probably confused with anthropic seething about "distillation attacks" coming from "fraud accounts". But that is not the law, that is just Anthropic being upset.
AFAIK this does not set a legal precedent as it has been settled and last summer finding is that Anthropic was wrong for "acquiring books illegally" not for training which is fair use.
With model distillation being so effective now nobody actually needs to pirate books to train their models. You can get an open-weight Chinese model and get all that. Or you can just buy the books or buy a library - there are many creative solutions here that aren't piracy and not going to cost you billions of dollars.
The moat right now seems to be the compute resources which might actually be worse for us common folk than a legal moat as we need compute for many more things that aren't LLMs too.
> and then sent to all its subscribers, things need to change
and then sell to all its subscribers, things need to change.
Fixed that for you.
Imagine being able to pay a fraction of your savings to download all Netflix shows and then sell 1 minute chunk of every media to your paid subscribers.
Exactly. This is a slap on the wrist. They need to either be banned from profiting from the egregious piracy, meaning charging money for anything trained on pirated works, or at least be forced to pay major royalties.
Regurgitating existing ideas is not copyright infringement. Reproducing works verbatim is and AI companies already implement guardrails to prevent that.
It depends. In music for example it's often question whether the artist has been exposed to the original work. In that spirit, small language models are less likely to infringe copyright.
This is exactly why something more advanced than copyright is needed to protect human creative endeavour against AI appropriation. Copyright is demonstrated here not to be up to the job but it doesn’t mean there isn’t regulation needed to give human creators rights and a reward for their contribution
What you are saying leads to "pulling the ladder behind you" effect on creativity. It's impossible to protect more than substantial similarity and still allow creativity to exist.
If a human makes something there should be broad protection for creativity, if a LLM generates something there should be extremely limited protection for creativity.
You should not be able to mass generate images in a particular artists style and claim it as fair use, even if a human making the same images would have protection.
But the human made the LLM. An LLM is categorically “I built a thing that built a thing” and if the output of that category has no protections then all automation and ‘machine at the final step’ is in trouble.
What about aleatory music (music left at least partially to chance)? Or Autechre - they have whole albums and live performances built on automation software. They built the logic and added randomization, necessarily removing themselves from the final output.
Is spin art not copyrightable? If I build a simple machine that spins paper, then do no more than drop paint on it, the result is not mine to copyright? I didn’t choose the output, I merely built the machine and the rest was created by pure chance. “But you chose the paint” - and if I didn’t? What if my art uses AI to perform sentiment analysis on the top news articles of the day and it drops colors matching the emotional tone of the news onto the spin art machine. I have no control over it and the output is machine generated, but is the result not just the final step of an entire process I created? Was the result of the creative idea not part of the creativity itself?
If I build an automated laboratory to test every combination of a problem space, is a resulting success not patentable? What if the problem is too large to permute, so I added a random selection process to it? I’m not even controlling what’s being tested, but if it finds success is that not my contribution? The machines did the work, the selection was random, there was no human in the loop; what then?
The internals of an LLM may be mysterious to some, but I assure you it’s just fixed automation with a random number generator sometimes tacked onto it, but randomization is optional too.
I built that LLM. I decided what text to input for training, I curated the information, I wrote the algorithm, I decided the layers and hyper-parameters, I decided the RLHF pairs to train, then I put a few drops of paint from my bottle of language into the automated machine. I decided and built every single step of the system, but that output is not part of my process? If I pipe the LLM text output to a paint dispenser hovering over paper, set to squeeze out drops based on syllables, would you protect my artwork then?
It would be interesting to know what the guardrails are. That would help me with understanding how I can use AI content. For instance I asked Claude to help me draw a diagram to represent a software engineering concept for a public presentation and then I had to stop and think: am I about to just reuse something from a Martin Fowler or Kent Beck book without attribution?
I think the only way to stop that is to put the responsible folks in prison permanently. Small criminals are being jailed permanently on repeated offence. I think big guns with a lot of money need to get much higher sentences by default. And no monetary way to avoid that. The whole prison system is kind of screwed up here. A leech system for lawyers and judges.
> There needs to be a royalty payment based on if the AI regurgitates existing ideas.
So, by that logic, you need to be paying every time you regurgitate any of my ideas. Or anyone else's. Copyright now protects abstractions and vibes. Substantial similarity test be damned. Nobody can write stories about wizard schools, the idea is taken.
Humans are not computers. Humans are not a service. In the end, all laws are made up rules and can absolutely be written to have different outcomes and restrictions based on if a human is doing something or if a program is doing it.
Real question, if an LLM shouldn't be able to remix someone's written work, why should a robot be able to build a chair that kinda looks like a chair a carpenter built that one time? The carpenter was a human, and humans are not a service.
You indeed need to pay someone if you take their copyrighted materials and regurgitate it. Ask DJ's and producers how they need to include royalties for samples used in their tracks.
There’s a difference between an abstract idea and the concrete thing. Regurgitating an idea is different than repeating the text verbatim. Ideas are protected by patents, not copyright.
So now that we have a magical paraphrasing machine, we can just run any copyrighted work through it to remove the copyright? Cool, I get a GPL version of Microsoft Office.
Maybe but your brain is not running 24/7 capable of outputting thousand if not millions of tokens per hour, all while having ingested nearly the entire internet.
If yours do that, maybe we can redefine what copyrighting and patenting means for humans
playing an MP3 at a venue, streaming it or distributing it is a copyrighted act because, despite not being a verbatim copy of the original material, it is capable of producing a nearly-verbatim version of that intellectual property well enough that most people won't be able to notice the difference.
similarly, as has been shown (by numerous publishers and authors), LLMs are capable of producing nearly-verbatim versions of the texts they have been trained on, to a well enough quality that most people won't be able to notice the difference.
the fact that an MP3 cannot "paraphrase" or "summarize" the audio data is not what makes it copyrighted, and neither does the ability of an LLM to "paraphrase" or "summarize" the textual data it's been trained on, make it any less intellectual property theft
the motivation for the audio case is the sense that the listener will not care whether the DJ plays an MP3 (they didn't pay for) or plays the original record (they would have paid for).
similarly for the lossily compressed text engine aka LLM's case, many people will not care whether they get this textual information paraphrased or nearly verbatim from an LLM trained on pirated books, or the original books.
the fact that an LLM also has the ability to paraphrase or summarize the pirated textual information it's been trained on, doesn't really matter if it's also capable of producing nearly verbatim copies of (parts of) those texts.
to underline this point even more, we know that MP3s (and more modern and much more efficient codecs like OPUS, after that) have been psycho-acoustically optimized to store exactly the least amount of data that will get "the point" of that music across to the listener, to the extent that they do not need the original recording any more. this is the stated goal of lossy compressed audio, after all. well, it also happens to be the (pretty much stated) goal of LLM companies, to store exactly the least amount of data that will get the point of that text to the reader. and it does tend to cause the readers to not really care about the original book any more.
having said all that, I don't mean to argue to lock it all up. I actually mean to argue that we should demand that Anthropic and Open AI release their weights data, and if anyone were to happen to break into them and steal that data, I would have exactly zero pity for that. because fair is fair.
> similarly, as has been shown (by numerous publishers and authors), LLMs are capable of producing nearly-verbatim versions of the texts they have been trained on, to a well enough quality that most people won't be able to notice the difference.
If that is true, you have a legal claim and can sue them. I doubt that’s true in the general case though.
The “does it hurt the original publisher” is a test for fair use BTW, just because you hurt the sales of someone doesn’t necessarily make it copyright infringement. That is only relevant if you try to defend using fair use (and it’s only part of the test that’s used to decide fair use).
It's a perfectly sensible interpretation of international copyright law.
I'm not sure if you're serious with the suggestion I could sue them. These are both US corporations, that justice system is pretty much in shambles in particular when it concerns corporations as big as these AI ones. You can dig your heels in the sand to defend that system, but you will also have to dig your head in the sand about why Sam Altman doesn't have a Disney "influenced" avatar, but one "inspired by" Studio Gibli.
And I'm not sure if you're familiar with the concept of "fair use" in the US as it "works" in practice, it's almost insulting, ask any music education youtuber.
Also even if it would work (which it very much doesn't), whether it "hurts the original publisher" is actually literally one of the criteria for considering something fair use or not. Look it up.
> Also even if it would work (which it very much doesn't), whether it "hurts the original publisher" is actually literally one of the criteria for considering something fair use or not. Look it up.
I don’t know why you’re repeating the stuff I just wrote like I didn’t. My point is that this is only relevant for the fair use defense and not copyright in general.
Here’s what I said:
> The “does it hurt the original publisher” is a test for fair use BTW, just because you hurt the sales of someone doesn’t necessarily make it copyright infringement. That is only relevant if you try to defend using fair use (and it’s only part of the test that’s used to decide fair use).
This isn't that out there in our current scenario. These models compress our collective thought and effort. Why not make these publicly owned, all profits distributed back to us?
It's not meant to do anything about LLMs. It addresses the procurement of training data. I'm glad the courts demonstrate some basic lucidity that sadly seems to have escaped tech discussion sites some time ago.
Ideas are not protected by copyright, nor are facts. You need to have a very specific and 'creative' / 'substantial' expression of an idea for copyright to apply.
The output of an LLM can be easily be such, but usually not.
That phrase is doing a lot of work. In the US, any writing is automatically protected by copyright. (This comment, for example.) Whether the author can claim infringement is a can of worms: legal costs, fair use … but your “very specific” phrasing makes it sound like there’s a prescription for exactly what is protected by copyright - there is not.
> Ideas are not protected by copyright.
The expression of the idea is, however. Same with facts. The fact that I live at a specific street address is not protected. My sentence construction explaining my specific street address is protected.
> The output of an LLM …
… is not protected, not matter its shape. The US Copyright Office has declared as much.
> You need to have a very specific and 'creative' / 'substantial' expression of an idea for copyright to apply. The output of an LLM can be easily be such, but usually not.
This is incomplete with current US law. You need the above (the typical copyright qualifiers) AND evidence of substantial human involvement in the creation.
Minimally directing an autonomous agent does not qualify.
Just to be clear, what you're referring to is the current US standard for whether a work is copywritable, not whether training on data and "regurgitating existing ideas" is fair-use. The latter is what the GP comment was about:
> There needs to be a royalty payment based on if the AI regurgitates existing ideas. That is probably the correct way to legislate this. If anything a human does can instantly be copied by an LLM, and then sent to all its subscribers, things need to change
Correct. I just wanted to clarify the statement that parent made, as it seems like lots of people have a misassumption about the copyrightability of autonomous in the United States.
Expect it will be clarified and/or changed by law given how much money is at stake, but the current state is what the current state is.
If I were developing key IP with agents, I'd be very careful to document my human contribution.
This settlement has basically nothing to do with LLMs.
At least not as far as the courts are concerned. Alsup ruled [0] that feeding a book into an LLM is transformative and counts as fair use. Especially when they purchased a physical copy of the book, scanned it, and destroyed the original.
But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they planned to eventually deleted them afterwards)
Where Anthropic ran into problems is that they put all their pirated books into a big central library (file on a server), and planned to keep those copies forever. Including copies they never actually fed into the LLM (a point that seriously worked against them).
Alsup ruled this central library of pirated books was copyright infringement. And it's this "pirated central library" that Anthropic are now paying a a 1.5B settlement for, nothing else.
The fact that the pirated books were also used to train LLMs is legally irrelevant. Though... I suspect a non AI company could have negotiated a significantly smaller settlement.
> At least not as far as the courts are concerned. Alsup ruled [0] that feeding a book into an LLM is transformative and counts as fair use. Especially when they purchased a physical copy of the book, scanned it, and destroyed the original.
> But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they eventually deleted them afterwards)
The way I understood it, was that essentially the entire case rested on if Anthropics use was "transformative" or not. And since they literally destroyed the books (not just delete files, which would be copied), that made it transformative.
Regardless if they deleted files or not, if nothing existing was transformed, it would have been illegal. But because of the destruction of k̶n̶o̶w̶l̶e̶d̶g̶e̶ physical property, this ended up being legal.
> And since they literally destroyed the books (not just delete files, which would be copied), that made it transformative.
You have to be careful, just because the judge points a factor out as notable, doesn't mean that factor was required.
The destruction of source books makes Anthropic's fair use argument [2] especially air tight, but it would be a mistake to assume that act was required, or is what made it transformative.
In the previous google books case [1] (which this case cites), google borrowed books from libraries, scanned them, then returned them. They were not destroyed, google didn't even keep the physical copy.
Yet Google Books was ruled fair use, because it was transformative.
My understanding comes from here, seems pretty clear to me but won't claim to be a lawyer of course:
> Ultimately, Judge William Alsup ruled that this destructive scanning operation qualified as fair use—but only because Anthropic had legally purchased the books first, destroyed each print copy after scanning, and kept the digital files internally rather than distributing them. The judge compared the process to “conserv[ing] space” through format conversion and found it transformative. Had Anthropic stuck to this approach from the beginning, it might have achieved the first legally sanctioned case of AI fair use. Instead, the company’s earlier piracy undermined its position.
Based on that I get the impression it's quite literally the destruction part that makes it transformative, without it, it wouldn't have been tranformative at all.
I've read through the order again. I can't find anywhere where Alsup says the destruction was required.
He cites three cases where a conversion from one format to another (without destruction of the previous version) was ruled to be fair use. Including scanning books with the google books case. (And referenced the Napster case, where a similar argument was rejected)
Then made the following comparison.
"Here, every purchased print copy was copied in order to save storage space and to enable
searchability as a digital copy. The print original was destroyed. One replaced the other. And,
there is no evidence that the new, digital copy was shown, shared, or sold outside the company.
This use was even more clearly transformative than those in Texaco, Google, and Sony
Betamax (where the number of copies went up by at least one), and, of course, more
transformative than those uses rejected in Napster (where the number went up by “millions” of
copies shared for free with others)."
So it wasn't transformative because of the destruction. The destruction only made it "even more clearly transformative" than those other cases.
Like, how can destruction be required if there were previous cases where it wasn't?
The key legal point is not that Anthropic destroyed the books, but the key fact was that Anthropic didn't distribute the scanned copies. Alsup keeps returning to this point:
"But what matters most is whether the format change exploits anything the Copyright Act reserves to the copyright owner. Anthropic already had purchased permanent library copies (print ones). It did not create new copies to share or sell outside"
"But again, the replacement copy here was kept in the central library, not distributed"
The conclusion of that section doesn't even mention the destruction at all.
arstechnica isn't exactly wrong, the quote also mentioned "and kept the digital files internally rather than distributing them". It just put way too much emphasis on the destruction, and not enough on the lack of distribution.
The other thing that arstechnica are missing:
Antropic didn't destroy the books because they thought it would strengthen their legal argument. They destroyed the because it's a lot cheaper and faster to scan books by ripping off their bindings and feeding the stacks of loose pages into a document scanner.
> So it wasn't transformative because of the destruction
I mean, the parts of "in order to save storage space" and "The print original was destroyed. One replaced the other." again makes it clear (to me at least) that the destruction is pretty much what sticks out here that makes it "more transformative" (whatever that means) than the previous cited cases.
But yeah, agree that also "didn't distribute the scanned copies" seems to have mattered a great deal, as well as the destruction part.
Yet countless families, including old folks were ruined during untold numbers of RIAA suits because "converting to save space" is not a permissable use.
They used to go around destroying lives by the thousands after Napster was creating because of the invalidity of that argument.
It is a crime to make a CD of your MP3s and vice versa, and you cannot convert your VHS to DVD.
A billionaire does it at scale, well then saving space via format conversion is a grand, while the peons still can see their lives destroyed but with it hidden via the CCB secret panel. Two tier American Justice on full display. Bankrupty and seizure or worse for thee and billions for he. Format conversion legalized only for oligarchs, and of course, no appeal so it will only be a binding precedent on that one rich guy and nobody else. Tribe on both sides, keeping special rights for themselves that are illegal for everybody else.
The RIAA (and the wider copyright industry) were careful to never bring a case that might rule on the issue of "ripping data from CDs and converting to digital".
What they did was bring a case against Naspter, which ruled that ripping data off CDs AND THEN sharing it to millions of people over the internet was infringement. Not because of the ripping, but because of the sharing. The RIAA then somehow managed to twist public discourse to interpet the ruling as "ripping CDs is illegal".
They were careful, because the Sony Betamax case had already ruled that recording TV of the airwaves was legal, which is already a weaker case than ripping CDs you own. They knew such a case would likely rule against them, and they found the ambiguity to be much more useful.
And later cases like the google books case, and this Anthropic one provide even more evidence that the courts would likely rule that ripping CDs was legal if such a case was ever bought. (Though, it really depends on what you do with the digital copy)
This is not backed up by any evidence. Ripping CDs was never illegal. The DMCA made the circumvention of an effective copyright protection mechanism illegal, which made ripping DVDs and Blu-rays a crime. But that's separate from copyright itself. The RIAA sued Napster users not because they were converting files, but because they were obtaining them from others without a license.
There is a lot of evidence. I lived through it. Every family with children and an internet connection or MP3 player was terrified of getting ruined suddenly via a letter. It was in the news every day about some other grandpa or single mother losing their house.
Ripping CDs was long illegal. Perhaps the Librarian of Congress made an exception. Now they hid everything behind a CCB that is like Arbitration so we will never know because they have hid almost aspects of societal justice about copyright and business labor behind arbitration style secrecy. The most useful courts are secret and now people believe there are no proceedings and they do not understand how much of our society was litigated and debated before.
Here is an article from 2008 Specifically explaining that ripping a CD to format convert for personal use is illegal and the RIAA and Sony BMG saying it merited suit but they had bigger fish to fry.
Mr. FISHER: That's right. So then, you have to ask yourself, why is the industry continuing to cling to that notion that there is no such legal right? (1)
"Bigger fish to fry" is not the same as "legal".
People selling software to easily convert VHS to hard drive were also punished. For decades they were very clear that format shifting was outlawed. But now that it supports centralizing power and creating a permanent class of info-priests to rule the society, they allow it for them.
Frankly, making all the justice system secret is why the media had to turn to personality cult nonsense for most reporting. All the great stories of the past were informed via the justice system activities. Since all the court stuff is secret now, all they had to talk about was Donald Trump.
"Discovery" provided the bulk of news facts before they secreted away all the justice system proceedings for liability, labor, negligence, medical care, copyright.
It used to be possible to know stuff about America and there was "evidence" all over the place. Now there is never any evidence for anything anywhere. That's Scalia's legacy thanks to Concepcion, absolutely gutting the ability of the society to use Hawthorne effects to discern legality and behavior.
This country used to have evidence for everything, and now a lack of evidence is so common that it is a trope level popular refrain.
You're confused about what was actually illegal and what the industry wanted you to believe was illegal. They didn't want to take anybody to court for actually ripping a CD because they didn't want to lose and have the precedent set like it was in Sony versus Betamax. This isn't "bigger fish to fry". This is "terrified of the precedent".
> The dispute arises from a suit the RIAA filed against a man in Arizona who bought CDs, copied them into his computer as MP3 files, and then put them into a shared folder that other people could access through Kazaa, a computer program for sharing music. He's being sued for that last part.
Then they state what they wish were true:
> But according to Marc Fisher, legal documents and some statements by industry officials make it clear that the industry regards the simple act of copying a CD onto your computer or your iPod as illegal.
But just because they wished it to be did not make it so. Trillion-dollar companies have provided end-users with software to rip CDs (including iTunes), and there's never been a court case over it.
They did not change the laws. The rich tribe guy ignored them and they let him off just like the Anthropic guy. They did not change the laws. The old businesses just folded and the new ones via unlicensed independent contractors made cottage industries out of small scale fraud and tax evasion.
How many Uber drivers can show their local business license for every town they pick people up in? How many have sales tax accounts for their state? Every uber driver without them should have been charged with the same crime as Al Capone.
Now they are trying to control the knowledge, and the vehicle driving, etc. via AI.
It is frankly a tribe takeover via mass criminal activity.
He should have been charged with tax evasion for every pickup in a place where he lacked a business license, but the tribe would never allow it. Compliance is only for the other guys.
Only the dumb local guy graduating high school trying to earn a living has to worry about legal compliance since the rich guys are too hard to prosecute.
They did not change laws. They refuse to enforce them and we are being taken over by the reincarnated legion of Al Capone as a result.
Nope, that's a little bit of sloppy writing on the part of Ars. I am not a lawyer, but I'll be happy to discuss the technicalities with anybody here. I'm fairly passionate about the technicalities of copyright.
The transformativeness of the use is independent of the destruction of the books. The destruction of the books allowed them to argue that they had not duplicated them, and was instrumental in the argument supporting the legality of scanning them. But that's entirely upstream of the way the data was leveraged, which is what is critical in the argument about the use being transformative.
It's easier to ask for forgiveness than permission, right?
It seems to be the modus operandi of corporations in general: they commit any kind of infringement they want and then later they go for a settlement with a value that's, of course, not too big for a company too big to fail.
In the meantime, the average person or company gets shafted.
In my opinion, we are one step away from AI companies capturing the entirety of copyright legislation.
You do indeed appear to have a valid point. Many "chosen" companies, like Uber for example, appear to have broken numerous laws. Legal action against many such companies comes suspiciously slowly, where they have already obtained massive profits and value, before the possibility of being shut down comes. Then, when they are finally pulled into court, they have all kinds of money for the best lawyers and have already paid the right politicians (and others).
When the legal judgements for wrongdoing are finally handed out, they often come across as just an inconvenience or kind of tax, which is easily handled in comparison to the profits they've already made. Yet, if average Joe or persons not considered as being of "the right type" were to do such actions, they quickly get the full book thrown at them. Often, the full measure of legal punishment, where their company and life is or about nearly over.
Paying this sort of fee in the first place is itself regulatory capture because only the big companies will be able to pay it. If they can pirate to make an LLM then so should us commoners be able to too.
As far as I'm concerned, the courts are wrong, and training on ill gotten copyrighted material is not fair use. Given the clear value of highly trained LLMs, the investment they have taken on, and the amount of disruption to the existing economy they stand to make, in a just world, the people who created the training data deserve some level of compensation. I think, in the US, they are very afraid of falling behind China, who doesn't give a shit about intellectual property, but that doesn't mean we aren't crossing an ethical boundary, acting like them.
Copyright is what stops someone from copy+pasting a book that took years to write, then selling it $1 cheaper than the original author on Amazon or whatever and making a margin 1 million percent higher than the original author.
Imagine a society without copyright… only physically intensive jobs could make money because everything else would be pirated, ripped-off or free. Thus, only those who are financially independent could afford to publish. Because the world really needs more rich class propaganda…
Right, it's about incentivising intellectual work. While I have big issues with the copyright system, like all the extensions lobbied for by Disney and friends, it did enable a lot of good work to happen.
How do we know that when we don't have a copy of the world without this regime? How much more and greater works could have been produced without such a repressive system?
A really successful work becomes part of the culture, and remixing, derivatives and other modes of integrating cultural artifacts are prohibited. Why should we allow corporations to own our culture?
Isn't most creative work synthesis rather than unique whole-cloth creation? Look at what happens with software when it is open sourced and allowed to be remixed freely. Are we better or worse off because of it?
There's nothing that prevents people from remixing things that are not copyrighted and create something amazing that others are interested in or of cultural value.
With open source, I should note, its remixing is in fact governed by copyright.
I think the idea is if "AI" can solve math proofs that humans haven't for a century then if "AI" freestyles stolen art and literature then it might create something as good if not better because of resources and processing power
There might be something to that logic but art and literature doesn't obey rules like math and copyright exists to protect creators
Yeah, I can't AB test against a world without copyright at all, but I think there's sufficient evidence to believe that a lot of stuff would never have gotten done without copyright to ensure it could be done gainfully.
The importance of striking a balance between incentivising creation and enriching culture was why the original copyright term was dramatically shorter. The modern term of owners life + 80 years or whatever it is, is clearly ridiculous. 20 years before entering public domain seems pretty reasonable.
There's unfortunately also some pressure against people using legitimate public domain works. E.g. youtubers getting copyright strikes for playing public domain music because it's too similar to a specific copyrighted recording.
Quite rich coming from a creator apparently, only certain creations are deemed worthy by you, seems like it invalidates your entire stance on the sanctity of art that all these pro-copyright people seem to hold.
Copyright doesn't actually stop me from pirating a book or an mp3 right now. Heck, I'll just download a book right now. Bam. Done. Some things are so difficult to keep from being pirated, such a photographs, that saying the copyright system protects photographers strikes me as a bit silly. It does protect some commercial photographers if a magazine wants to sell their photo sometimes, but that's a very very small slice of all the photos in copyright that are being shared online right now.
Also there are other systems that might protect an author's financials. Off the top of my head I imagine you could do a netflix model where every citizen pays some taxes to consume intellectual property like a utility. Then the goverment finds a way to measure what is being consumed and gives each author a share based on the rate of consumption. In fact the "intellegence is a ultility" ramblings of Sam Altmen sort-of point in this direction. But that's just one idea thought up early in the morning when its too hot to sleep properly. I'm sure there are many others.
> It does protect some commercial photographers if a magazine wants to sell their photo sometimes, but that's a very very small slice of all the photos in copyright that are being shared online right now.
That is a very small slice thanks to copyrights. Without copyrights then corporations stealing from the small guy like this would be the majority of it.
How many times were hugely popular books rejected before a publisher decided they were worthy?
Copyright far more protects the wealthy than the good. They don't need to sell your book, they just need to own the book that people are buying right now. Giving your book a chance to sell would dectract from those sales.
If there were no copyright anyone trying to sell the book $1 cheaper would be undercut by someone selling $1 cheaper them them, and so on. The financial incentive to do that goes away. People then choose to distribute based on different incentives, like the fact that they have seen something worthy that others should see. We have almost completely lost that today because the financial incentive doesn't care what it is as long as you buy it. That might lead to a world dominated by an optimisation for whatever it takes to get you engaged, or worse, addicted. That world might really suck.
There needs to be a way to support the creation of art. Copyright lets a few corporations decide the subset of available art is seen enough and available to pay for (in the hope that maybe some of the patment gets to the creator). It is not a system that works in the modern world.
Great, that way is called copyright. The author has the right to control who has the rights to distribute their work, and can require compensation in exchange for that right; what economists refer to as "selling".
If money is the only incentive, then it's not a product of artistic work.
Also current copyright laws only exists to fulfill the constitutional mandate to promote the progress of science and useful arts. There are a lot of alternative ways to fulfill that mandate that don't include a lot of the baggage we have presently in copyright law which is now slowing down progress.
>If money is the only incentive, then it's not a product of artistic work.
This is just bullshit and no one said it's the only incentive.
> There are a lot of alternative ways to fulfill that mandate that don't include a lot of the baggage we have presently in copyright law which is now slowing down progress.
You're speaking to the generation of pirates. What? Suddenly everyone is hanging up their high seas hat to capture the virtue signals of current sentiment?
It's funny because copyright only benefits the rich now. Record labels hold all the copyright to songs, same with publishers for books, Disney made sure it lasts over a hundred years. The days of copyright being held by individuals in any real sense is long gone.
Never really understood how libertarians expect to have someone making guns for their fiefdoms when there is no one to enforce property rights for said gun elements and manufactories.
Libertarianism is not a philosophy. It's selfishness taken to extremes and trying to find ways to justify it at a societal level. The only reason we're the top species is because we're ultra social and have culture, which is inherently a social trait (don't eat those red berries, they're poisonous). Libertarianism want all the benefits of working together with no actual thought into how that working together happens in real life, including punishment for bad behavior.
I guess they presume it requires on the good will of everyone to live peacefully without violating (non-existent) property laws over, say, robbing you in your sleep.
I would be fine with abandoning copyright ... If it is done for everyone equally, and not just tech giants and VC money businesses get a free pass, while everyone else still has to follow the copyright laws. Lets go ahead and usher in an age of free information and experiencing all forms of human expression for everyone. But lets also come up with a way, to compensate our creative minds and our educators and artists. How about that UBI? We stand much to gain as humanity.
This. Copyright is a flawed system. There can be alternatives that allow more than 1 player to play and not create monopolies.
For example. I invent a new method of power washing. I start a power washing business using new tech. I file the tech for patent and copyright-equivalent use. This is then made available to other power wash companies that wish to use the tech and be certified in it so long as a small portion of their revenue goes back to the inventor for a set amount per volume, or something similar of a metric that has a cutoff after a point.
This will breed new industries, create new jobs, introduce new innovations, and allow the markets to move on from being strangled by one giant corporation.
So will you owe life long compensation for all the knowledge you got from books too? How about all the pirated books, music, movies, etc you consumed? When will you set up a life long payment plan to corporations that own these rights, because I have a bridge to sell you if you think any of this settlement will go to any of the people who created anything.
I’m guessing you have some kind of imagined idea of some small author being compensated handsomely for his book and all future earnings that could have come from it. Reality though is that between the attorneys that will run away with some high triple digit millions and the corporations that own the rights to the subject works, there will be measly “checks” for any actual person that created anything, i.e., an artist or author.
In an odd way, this whole case is really just “capitalism” cannibalizing itself, i.e., publishers greedily and also in a terrified manner trying to steal away as much capital from the technological shift to AI as possible in order to either create a buffer or fund their transformation to adapt to what AI means to the very nature of writing itself, let alone publishing.
I suspect human writing could survive, but I don’t see any room for publishers.
I think people forget that laws are perfectly capable of carving out exceptions, leaving purposeful ambiguity, expressing intent, etc. Yes, humans can have special rules, and very obviously should since laws exist to improve human lives.
I am actually not settled on either side of the matter and I have not forgotten that, but I think what we are really looking at is a rather more complex matter than people want to make it out to be. We are holding several but at the very least contradictory positions and they are incompatible.
Either it is theft or not theft. Why would you stealing from me deserve an exception, but when a group of people in an organization you may refer to as a corporation steal from me, you want them to pay me? Do as I say, not as I do or something like that?
Of course exceptions can be carved out, but they cannot be just, inherently. The problem is that we have allowed our ruling maniacs to create a fiction that organizations are people, which not only have more rights, and less responsibilities, and even less consequences/penalties; but also confers upon the individuals that make up the corporate person rather extreme super powers like being able to commit crimes up to outright murder, and there not only are effectively zero consequences for or to them but in most cases today they immensely profit from it and then shield that money from the victims seeking justice.
The underlying issue, why I am not settled on this matter, is that it is inherently contradictory because the facts and underlying assumptions are all so distorted and perverted that there is no good answer to be had and it's really just a matter of rule of power, feigning rule of law.
> So will you owe life long compensation for all the knowledge you got from books too?
You're just falling into the trap of anthropomorphizing the phrase "training" in the context of LLMs, which is not the same things as what humans do. There is no evidence they are the same thing and there is nothing to support the notion that what an LLM does when it "trains" on a book is equivalent to a human reading it.
Whatever "training" is, if you can't persuade the machine to spit substantially the same text back out verbatim, it's clearly not something that falls under copyright law either, because there's no copy.
Yes, for some texts that's possible. But for the vast majority, it is not.
Can you cite any information on this not being possible for the vast majority?
Or is it simply that the correct prompt hasn't been written for all possible cases?
I also fail to see the difference if logic/harnessing is added around a vector database that can output the complete corpus, but simply is instructed not to.
It very clearly is still compressing the information into the vector weights, and then recovering that information, thus the information is encoded.
Why is a vector database somehow completely different from maintaining a library of the text itself?
Is that relevant? I can use a lossy compression algorithm such that the original could never be recovered from the image I've produced, but that derived image would surely be under copyright.
LLMs are obviously capable of producing "exact" phrases as well. Ask it to give you famous quotes, it can do it. Ask it to read a paper for you and cite it, it can do it.
Because that's legally the case? I don't understand the question. Using a lossy compression algorithm on an image does not remove its copyright protection.
You are asking to prove a negative. But even assuming that the model is capable of returning every bit of its training data verbatim (a mathematical impossibility) that would not be enough as mere capability is insufficient here. If capability alone were the standard any library that also has a photocopier / scanner would be in violation.
To prove distribution of copyrighted materials it would have to be practical and actually used in the wild by people to circumvent copyright and generate copies of those works. Again, I can't prove a negative, but that isn't the standard, and nobody has shown a practical exploit here.
There was a paper a while back where (from memory) they managed to coax 75% of the original text of some internet-popular books out of an LLM. Harry Potter, 1984, etc. That's why I said it was possible for some texts.
My assumption is that multiple copies in the training data "wear a deeper groove". I believe those are infringing, and should be dealt with on a case-by-case basis. But the vast majority of text doesn't wear that groove.
Nah, I can't prove a negative. But Common Crawl is 12 petabytes and is not the largest part of what these models get trained on. DeepSeek v4 Pro is, what, 865GB?
That's one hell of a compression ratio, if it can do what you claim.
The New York Times lawsuit is resting on the point that large chunks of undigested articles can be vomited out. OpenAI tried to have the lawsuit thrown out but the courts permitted it to continue.
The Times... alleged that OpenAI's ChatGPT and Microsoft's Copilot had produced near-verbatim replicas of copyrighted articles, that the chatbots generated hallucinated content falsely attributed to the Times, ...
Don't get too hung up on the preciseness of the copy - the courts won't. I doubt that spitting out an existing article with a few adjectives changed would be considered transformative.
It's possible. Would be interesting to see their evidence, and to know whether they can reproduce it for arbitrary articles, not just ones that have been endlessly republished on the net.
This reliably inevitable rationalization comes up in every thread it seems, and its ultimate goal is to humanize AI. This is what the big guys want us peons to believe and it works so well, I have even been lectured by an AI for being rude, the implication was that I was logged in and it would be a shame if anything happened to my account.
Quit trying to make AIs human, people who are trying to make AI human keep forgetting that humanized AI's have only the morals relevant to their mission, there is no profit in humanizing AI's because if we continue on this track of humanizing AI's, we being stupid humans will grant them civil rights expecting these new AI's with rights will somehow respect our rights and thats a fundamental misunderstanding of how AI'S actually work.
> As far as I'm concerned, the courts are wrong, and training on ill gotten copyrighted material is not fair use.
It’s important to remember that a court’s job is to apply law to a situation. When a court gets something wrong it’s a misinterpretation of the law and will, by definition, be overturnable on appeal. I suspect that your objection isn’t that the court is wrong, it’s that the law is wrong.
It's not a settled area of law and there is a SDNY judge that has a completely different application of the fair use analysis in the same exact context and came to a completely different conclusion (that it is not fair use).
Sorry, I'm thinking of Kadrey, where the court rejected Anthropic's "training" argument and provided an explanation as to how author litigants should demonstrate market harm in order to succeed on a fair use analysis, a factor that Alsup did not effectively weigh.
I suspect the market harm angle is not going to work out either based on the one study I know of on the topic: https://www.nber.org/papers/w34777
> We document a tripling in the number of new books coming to market between late 2022 and late 2025 that mirrors the use of AI that we detect in new books. The effects of this influx on consumer welfare depend on the quality of the additional books. The average quality of new books has fallen with the LLM-induced influx, and books with detected AI are substantially worse than human-authored books, so that much of the new work is of little value to consumers. Still, the LLM influx has delivered some books in the middle range of the usage/quality distribution, and the LLM-era entry process delivered seven percent more consumer surplus from books than the pre-LLM process in 2025.
...
Moreover, the arrival of LLMs does not appear to have displaced activity by incumbent authors. Despite the controversy surrounding
LLMs, their effect on book consumers – like other cost-reducing technological changes in the cultural industries – is positive. However, because the new books are mostly of low quality, the effects are modest</i>
So not only are existing authors unharmed (because most of the new competition is slop) there is even a small improvement for consumers.
Yes, ultimately the problem is that the law is vague or inadequate. The courts have their definitions of fair use, which are their best efforts at interpreting the law, and I have mine, which is different.
William Roper: "So, now you give the Devil the benefit of law!"
Sir Thomas More: "Yes! What would you do? Cut a great road through the law to get after the Devil?"
William Roper: "Yes, I’d cut down every law in England to do that!"
Sir Thomas More: "Oh? And when the last law was down, and the Devil turned ’round on you, where would you hide, Roper, the laws all being flat? This country is planted thick with laws, from coast to coast, Man’s laws, not God’s! And if you cut them down, and you’re just the man to do it, do you really think you could stand upright in the winds that would blow then? Yes, I’d give the Devil benefit of law, for my own safety’s sake!"
This is why the idea of being "Vogelfrei" or "lawless" was honestly a terrifying concept in the middle ages. They are neither bound by law, nor protected by law.
A lawless man can be struck down with force without persecution by law, because they are lawless.
"I haven't loaded an advertisement in 20 years, I have 6TB of movies, 2TB of music, and seemingly endless file trees of mangas, all acquired for free over the years. Now having not said that, I beg you enforce copyright on these AI labs, so I can get a cut of their revenue for my years of writing well researched comments on the internet"
The internet, in true internet fashion, still has the general logic level of a 15 year old.
IMO, using copyrighted works to train models should only be "fair use", if the models are then released as (at least) open weight, so that the public can benefit from it. (Although as noted by a sibling, this would require a law change, not action by the court).
Adobe’s ereaders had a disclaimer that their books cannot be read aloud. There’s clearly precedent that this sort of transformation was disallowed by publishers at the time. Interestingly, at least the audiobook of the latest dungeon crawler Carl has a disclaimer that it can’t be used to train AI
That's a good ruling, because otherwise only the big companies can afford to pay for enough content to make an LLM (say goodbye to open weight or research LLMs). Having a fee like this is actually a form of regulatory capture.
> But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they planned to eventually deleted them afterwards)
The court says otherwise.
> Such piracy of otherwise available copies is inherently, irredeemably
infringing even if the pirated copies are immediately used for the transformative use and immediately discarded.
Then it says it doesn't need to decide on that basis because they kept it not just for training LLMs, but also for building a central library. Which seems a bit ridiculous, because the sole purpose of the central library is to train LLMs.
Who would have thought, that this is the way, which we take to arrive at the burning books stage again? They neatly line up with historical perpetrators in that regard.
>But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they planned to eventually deleted them afterwards)
You're not. Even if training is fair use, it doesn't mean you can steal copies to train the model. It just means the training itself isn't an infringement (in Alsup's opinion). Stealing the copies of the books was an infringement and that's exactly the liability that Anthropic settled.
Hey I just came up with this idea, I'm going to feed copyrighted books into my LLM that remembers them verbatim, and then people pay me to ask the LLM for complete copies of a book.
Wait, no, not verbatim. It transforms upper case into lower case and vice versa.
The correct way to legislate this is to abolish copyright. It is strictly a negative force. Nobody makes art because of copyright, only in spite of it.
Commercial enterprises stand on the shoulders of lax copyright laws. For example, foundational Disney works would have been illegal for them to make under the copyright laws they have since purchased.
What copyright law helps them? They have the least to worry about copyright as even if someone copies the movie script or something it's not like their views will be gone because of that.
I am not against trademark. e.g. Disney has a right on who can sell Mickey Mouse figurine, or Marvel has right over Iron man character and franchise.
People make art to also get recognized for that art. Otherwise they would keep that art secret at home.
Without copyright, anyone can copy the art and call it their own. What is then the incentive for the creator to share the art, if there is neither monetory gain and nor fame. And worse than them being recognized, they might even get accused of copying their own art if someone else became famous due to a copy.
Being "accused" of copying literally does not matter if copyright didn't exist. This framing is only an issue under copyright.
> Society would miss out a lot.
Society actively misses out a lot. We could have had tons of derivative art that has been buried for the sake of propping up companies. We could have had Aaron Swartz. Abolish copyright.
> There are lots of famous artists that are older than copyright.
There are lots of famous artists that made art before color image capture and reproduction (1930s-60s) and digital image capture and reproduction (90s-00s).
Copyrightless artist fame and economic viability is enabled by a lack of widely-accessible, cheap reproductive methods.
> You mean like right now? Here is A24 claiming copyright for Backrooms related media...
Turns out when you outsource copyright policing to minimum wage folks, ambiguity goes out the window.
> Society actively misses out a lot. We could have had tons of derivative art that has been buried for the sake of propping up companies. We could have had Aaron Swartz. Abolish copyright.
The problem with absolutist arguments is that they ignore inconvenient facts.
At a time when creative and art economics is under siege, how would copyrightless art make enough money for the creators?
The fact of the acquisition of large swaths of copyright rights by large corporations does not negate the fact that artists need food (and ideally, a place to live and a way to provide for their family).
There are positions that might enable that (Hey, what if we banned the assignment of copyright to corporations? Human only? Original creator only?), but none of them are stripping all rights from intellectual property.
Many of the most valuable paintings in the world are out of copyright. It has not diminished their value, because people still value originality even if it's not enforced by the law.
> And worse than them being recognized, they might even get accused of copying their own art if someone else became famous due to a copy.
This happens now all the time, and the winner is determined by who can afford the best lawyers.
> There needs to be a royalty payment based on if the AI regurgitates existing ideas
This does not do enough to fix the root problem.
People who live right now, who happen to have written or produced anything that AI works with, build on the back of humanities combined knowledge, will become outsized beneficiaries of AI, with the AI wave offering new ways of monetizing their work – while everyone who has not, won't be.
It's simply not good enough. We have to make sure people broadly benefit first and foremost.
It's easy to beat up on OpenAI and Anthropic, because they have lots of money and knowingly broke the law, but writing the book was onetime work too. Do we really want to turn everything into recurring revenue stream to skim of? How would that even work for an open weights model? Would you say the same about a human educating themselves from a book? The answer has to be more than pearl clutching for poor starving little authors (and the not so poor class action lawyers).
I agree. If I pirate a book and share it on the web, and I get busted for doing so, and subsequently pay a fine, I don't get to KEEP sharing it on the web.
Now, if I license the book, I might be able to come to an agreement with the author/publisher whereby I can share some of it.
The post specifically proposes royalties for ideas from books, not royalties for the books themselves. You would absolutely still be able to share ideas you learned from the books you pirated in that situation. It'd be insanely draconian if you couldn't.
(Then again, US copyright law often is insanely draconian.)
I'd rather see AI studios do the same as the film industry, pay a one-time up front cost per major model (or major.minor?) depending on how they contract it out. This also allows smaller startups to license books for less than a larger frontier studio would. In theory and hopefully, the pricing would not be too insane per book, you want them to rent more books and spend more, not go back to pirating right?
The big deal for publishers and authors is the payout per eligible title is $3k. For a traditional publishing contract involving one author, the amount will be split down the middle.
The other thing which caught my eye is the judge slashed the class counsel's fee by half, from 12.5% ($187.5m) to 6.8% ($101m). The class counsel's unreimbursed litigation expenses were $2.6m.
The three class representatives get just $15k each.
Realtors are capped in the percentage they can take for selling properties. Brokers and financial advisers are capped in their fees. My presumption is that the only reason this very standard and reasonable regulatory pattern doesn’t affect lawyers is because they tend to be the ones writing and enforcing the regulations in the first place.
There is no legal maximum for the percentage real estate agents can take, in the US. Rates are also not fixed, by law, and are required to be negotiable. There's a general standard for rates (typically 5-6%, split between agents/brokers), but there's nothing stopping them from setting it to 99%, other than the fact that people won't pay it.
Source: was a licensed real estate agent for a long time.
No. "Realtor" is a trademarked term for a member of the National Association of Realtors. Real estate agents are licensed by state governments, but prices for real estate agents are not legislated by state governments.
If this is the 2024 settlement that you are referring to, it did not say anything about the price a Realtor can charge:
>The cooperative compensation rule has been eliminated as a result of the settlement. Seller's agents are no longer required to offer compensation to buyer's agents when listing a home for sale on a Realtor-owned multiple listing service. In addition, Realtors acting as buyer's agents must enter into contracts with buyers before touring any homes, allowing buyers to negotiate how much they will pay their buyer's agent.
I'm aware that prices for agents are not legislated by state governments, but prior to 2024, as a realtor that was a member of your state's association, you were standardized at a 6% rate. Because that's what your association standardized. That's partially what the suit was about. What percentage of real estate transactions are done with agents who aren't members of XAR?
> pattern doesn’t affect lawyers is because they tend to be the ones writing ... the regulations in the first place
In my state, the people hired to write new bills (if passed becoming statute) have to pass their JD (law degree). Legislators only get to request new bills, they can't hand a proposal (which may have been written by a lobbying agency like ALEC or Heritage) into the system.
I've seen some YouTubes where lawyers were complaining about high bill rates and showing actual bills. One large firm billed the senior lawyers at $2500/hour and even the paralegals were billed at $600/hr.
Associates are billing over $1000/hr and can bill 12 hours a day easily enough. If you have 5 associates billing an average of 50 hours a week then in a month thats a million dollar bill just for the junior lawyers.
Alsup is an interesting judge. He has handled several important tech cases, such as Oracle v Google, and Waymo v Uber.
He's also a longtime hobbyist programmer working in BASIC, much of it in support of his ham radio hobby. Screenshots of his shortwave propagation prediction program here [1].
Yes, that is interesting. It sounds like he was aware of the theoretical possibility of a book being regurgitated verbatim. Do you know if he was aware it had been done? https://news.ycombinator.com/item?id=49000742
If he was not aware, I wonder if he still would have described the process as "exceedingly transformative" had he been aware.
Note that they're testing for 100-word passages. This is a level of memorization that avid readers can credibly also reach.
Note also that Sonnet 3.7 had to be jailbroken.
Note also that they got high memorization for a few books that were widely quoted. The books in question can probably also be "retrieved" by putting phrase prefixes into Google, which is probably why Sonnet 3.7 knows them with the precision of a fanboy. Material being widely repeated in the training set is a well-known cause of memorization.
No "avid reader" could recall anywhere near that much text. That takes dedicated effort to commit to memory. Copyright was never meant to stop people copying books anyway, it was meant to stop machines (ie. printing presses) copying them.
Edit: Apologies, I misread it as "100 pages". My point about copyright still stands, though.
I mean, the debate would then turn on whether publishing the passages and publishing the model is the same sort of thing. I think there's mainly two views: "we know the passages are in there, so publishing the model is publishing the passages is copyright violation", and "nothing happens until you go through considerable effort to elicit the passages, so the user is committing copyright violation using the model as a tool."
Personally I think our legal system is just not set up for a world where we can download mindstates in numeric form. Would a sufficiently detailed recording of my brain violate copyright? If simulated, it could certainly be elicited to commit violations.
edit: At any rate, Anthropic are not publishing the Sonnet 3.7 weights.
So, continuing to profit--forever--from someone's else work, at scale, without their prior consent, is fair use?
It's funny that crimes can be settled in cash. IOW, everything has a price; and the price is always right. Settlement ought to be the euphemism for blood money.
In addition to the settlement, what I'd consider fair is to have these companies pay royalties in perpetuity. Of course, that's not tractable.
Yeah, I feel like penalties here should be something like 10% of revenue in perpetuity. Then companies might think twice about asking forgiveness instead of permission.
You assume your premise. But plenty of "intellectual work" is already done without legal cover. It just typically attracts normal profits, rather than super-normal rent-seeking ones.
Rents are just amounts beyond what's needed to cause the thing to exist. At the point of copying something, it already exists.
So payments for the right to do so aren't payments required to bring anything new into existence at that point, save for the legal fiction.
Now you might argue that the future copy-licencing rents are necessary to bring the _original_ creation into being. But that doesn't make them _not rents_.
But I would say that's the second assumption you're baking in here.
As in, we live in a world where e.g. the movie Toy Story exists. Now, certainly Toy Story does provide some good or value to the world. But I don't think you can assume such things provide more value than e.g. open science, free transformation of works, etc.
I get that people enjoy our current IP culture but saying certain things wouldn't exist in an IP-free world is just an argument from consequences that doesn't even really compare consequences between the two.
Intellectual property protections have a finite lifespan to begin with, and as a tenet of Western Civilization are barely 300 years old.
For works published before copyright laws existed or after property protections expire, anyone should be able to use it for anything forever without consent.
Intellectual work still manages to get funded in this 'insane world' - although given the classical artist/patron system has given way to state-based grants and a select capitalisation of Art post-Warhol, the concept of Universal Basic Income tends to be promoted the desired successor.
Speaking of insane worlds, how does the concept of the Public Domain work in yours?
Are you really saying that because IP protections are (and should be, I agree) time-limited, they're not doing anything in the first place? This is patently ridiculous.
> Speaking of insane worlds, how does the concept of the Public Domain work in yours?
Can you elaborate on what you're asking? I don't understand your question.
Your contention was that it was an insane position that anyone should be able to use a piece of literature or a song for anything forever without consent. I simply highlighted the absurdity of that based on the fact that:
1. Copyright protections as a concept are an incredibly modern phenomenon, mostly limited in practice to Western Capitalist Democracies.
2. Outside of a short monetisable window (albeit one extended and irrevocably marred by Disney/Sonny Bono) your 'insane' hypothesis is in fact the status quo
3. Much intellectual work is published into the Public Domain, and all copyrighted work eventually ends up in the Public Domain. Your position appears to presuppose a world without such an entity.
As to what copyright actually achieves? It's mostly a mechanism by which the media gatekeepers and owners of capital use legislative and social imbalance of power to deny artist the rights and royalties for mechanical reproduction and otherwise impose financial serfdom.
This is achieved mainly by Copyright Enclosure, whereby musicians are typically pressured or contractually obligated to surrender their master recordings and intellectual property, and by contractual clauses like Controlled Composition Clauses, whereby Labels reduce the mechanical royalties they pay to artists who write their own songs, often paying below the standard statutory rate.
Lots of ways. Selling author signings, talks, authorized copies, subscriptions/merch, sponsorships. How do newfangled "content creators" fund their work? We already live in this world.
Maybe not all creative works are deserving of monopoly profits just by sitting on the ass in any case, and should stand on their own merits by producing downstream value that can be sold for whatever they can be sold for, by whoever puts in the work to deliver the value to the end user in a competitive manner. You know, open markets.
Attribution I can see. Consent or payment beyond market value, why? Just because you put in a billion hours to make a shitty $1 value output I should pay you a billion hours worth of labor?
Most of these turn intellectual work into that of indie musicians, or outright beggars. You are stepping dangerously close to stripping people rights in favour of giant AI companies.
> How do newfangled "content creators" fund their work? We already live in this world.
"Content creators" heavily rely on IP protections. You could always try taking some youtube videos with 100M views, altering them a bit an using them as your own and seeing how that goes down. Do let me know!
> Maybe not all creative works are deserving of monopoly profits just by sitting on the ass in any case, and should stand on their own merits by producing downstream value that can be sold for whatever they can be sold for, by whoever puts in the work to deliver the value to the end user in a competitive manner. You know, open markets.
And open market is not one where I can say that you are just sitting on your ass, so I'll take your stuff and sell it.
> Attribution I can see. Consent or payment beyond market value, why?
Because it's my stuff of course! And why should one even pay market value in your world? Why not always 0?
> Just because you put in a billion hours to make a shitty $1 value output I should pay you a billion hours worth of labor?
What in the world are you on about? If the price someone puts on their IP seems too high to you, you should not pay that price. We wholeheartedly agree. Where we disagree is where you from this conclude that you can just choose the price yourself and take it anyway!
> So, continuing to profit--forever--from someone's else work, at scale, without their prior consent, is fair use?
No, that's what they got in trouble for - a lack of consent.
If the author consents, it would have been fine. If they bought the books, then it is fine. Digitisation through destruction, like most book scanning systems. As long as the original work is destroyed during the process, and you actually paid for it, then it is fair use.
If it regurgitates, then the author can sue you again. So you are incentivised to make damn sure it doesn't. That's not covered by fair use.
Its only if the original cannot be accessed anymore, and you paid to get the original. Both must be true, for fair use to hold.
That’s what the got a _slap on the wrist for_. 1.5 Billion of a payout to effectively cement themselves as one of the only orgs that can ever create one of these models because the ladder is pulled up behind them.
Format shifting has a DMCA carveout. It is 100% allowed. Since around 2000, the rule has permitted it for:
> Literary works, including computer programs and databases, protected by access control mechanisms that fail to permit access because of malfunction, damage, or obsoleteness.
DRM being covered under other laws, and being gross, still applies. And still applies to industry giants, too. Which is why most who do this, like Google, actually buy physical copies and scan it destructively, so they don't have to deal with it.
Yes, it's legal to format-shift DRM protected media, but it is not lawful. Someone has to break the law to provide me with a decryption tool.
If you asked the right politician when all these rules were being written, the intent was that each person who needs to format-shift their media would independently write their own decryption tools, use them for lawful purposes only, and then dutifully delete them the moment they were no longer needed. This is, of course, laughable.
Of course, if Anthropic was, say, buying and decrypting Kindle books TODAY; they probably could get Claude to vibe-code a DRM decryption tool[0]. That would actually be within the bounds of this asinine law. If Anthropic started off by doing this, however, they probably would have just used a decryption tool found on the Internet, and that would have invited different legal challenges. Like, is it legal to use an unlawful tool to accomplish something otherwise legally protected? The courts so far have been very hostile to ANY attempt to tie the anticircumvention provisions of the DMCA to fair use. They could easily say "No, you only get to format shift with your own tools".
[0] Related note: I really wish I had Mythos access, just so I could jailbreak my iPad on modern iPadOS. No other reason.
I dunno, ever used a thing you learned from a textbook in your job? Did you have to continue paying for the copy of that knowledge speed on your brain? No, because that's not what copyright is about.
Learning from and building on previous work is civilization. Copyright maximalism is a plague.
Yes but does the computer actually learn? Is the computer a person that read a book and remembered, a part and used that to create a novel idea or is it just cioy pasting the answer and then reselling that
A computer not, but a LLM model does learn. The original text in no way exists 'in the model', and the model does not copy-paste it and resell the original text. Not the same, but reasonably similar to a human.
We may need some new legislation. An LLM is not a person, but its also not just a storage solution.
> I dunno, ever used a thing you learned from a textbook in your job? Did you have to continue paying for the copy of that knowledge speed on your brain?
Those regulations and principles are for humans.
Either the major LLMs are software tools deployed by ostensibly-profit-seeking companies, and regulations based on the notion that "making humans pay to make use of the things they've learned is profoundly antisocial" don't apply, or the LLM companies have a bigass swarm of unpaid -er- "servants", and labor laws and other human rights regulations do apply.
> I dunno, ever used a thing you learned from a textbook in your job? Did you have to continue paying for the copy of that knowledge speed on your brain?
They are, presumably, human. We can perfectly well say that humans have certain rights without needing to give machines those same rights.
For example, we've more or less all agreed that it's fine for a human to watch a movie and enjoy the memories forever, and be inspired by it forever. But we've also more or less all agreed that that doesn't mean that a human can use a machine to record that movie and keep it forever.
> Learning from and building on previous work is civilization. Copyright maximalism is a plague.
The debate has existed for several generations at this point. You may disagree with the mainstream opinion, but it's disingenuous to frame it as "copyright maximalism".
Well, as not every book is a textbook, I'd say quite a lot of what I've read never went to any kind of knowledge in my head at all. But I reckon the author still deserves to eat.
You too deserve to eat. Does it mean we're all supposed to pay you in perpetuity for that HN comment you just posted, because we read it and it's encoded in our brains now?
No... But as we're discussing people who tried to pay nothing... Maybe they should have just bought the books in the first place. And comparing my single sentence on HN, to the hundreds of millions ingested, does suggest that maybe scale changes something.
Like most things said on social media, not being copyrightable.
> So, continuing to profit--forever--from someone's else work, at scale, without their prior consent, is fair use?
Yeah, you are right. Have you been paying your dues to the authors of your math books in first 4 grades? I think 15% of your wages as an engineer would suffice. These kids continue to profit for years, and they are so many. Gotta pot a stop to that IP theft.
In my mistake I thought copyright was about copying rights, not paying for using ideas themselves. If just being downstream from a copyrighted work is infringement even without substantial similarity, then it's more like patents that expire in lifetime + a million years.
Because the process is the same. When you read a book you don‘t save it as a brain file, you form memories from it. Some people can recall verbatim bits here and there but I have never met someone regurgitating a book word for word. And I‘m pretty sure I can not ask ChatGPT to output the first chapter of Moby Dick word for word. I think that would be, rightfully, considered copyright infringement.
Why is it different other than, "just cause?" No one seems to have actual reasoning to back it up while it feels very similar the other way around, is human brains and neural nets (notwithstanding that they're both called neurons) seem to learn similarly and can act on similar classes of problems like language and mathematics.
What's the difference between humans and LLMs? The difference is that the creators of laws are humans, and the purpose of laws is for humans. Authors write books with the expectation that they will be read by humans, and copyright law was written with unstated assumptions, like the fact that books have an effect on a person's mind after reading it.
IMO the spirit of the law would prohibit LLMs from training, and the letter of the law leaves room for that only because nobody thought to write down "books are for people to read".
Laws are not only for humans, there are laws for bots as well, like anti spam, but that's besides the point because in reality behind LLMs there are humans and so humans still control them, therefore laws target them too, and now it looks like humans using LLMs to train via ingestion of books is deemed fair use.
He wasn’t just some famous redditor. Aaron Swartz helped create Reddit and invented RSS.
If I put my conspiracy theory hat one and I always get piled on for this theory in other online communities but I think it could be possible. The theory is I think Aaron found some very dark stuff while exploring the MIT private networks, things that he was not supposed to see and could be very damaging to a lot people if they were exposed.
The infamous Jeffery Epstein was donating a lot of money to MIT and its Media Labs. I think there is a much deeper story at play that the mainstream narrative is hiding with a “suicide”.
If he found fucked up stuff regarding crimes way worse than his on university networks don't you think he could have gotten out of the sentencing entirely by cooperating against them
Epstein was relatively restrained even in his personal email, I doubt he was using MIT administrated systems to facilitate a pedo ring
It's not that simple. If he did find dark stuff, and he got caught - he will have both the state murderers and the filth on his case. That's the best case scenario assuming the state isn't on the side of the filth by which case he was cooked no matter what he did. Filth does not even have come after you. They can just send you a family portrait and you'd know the hole is deep and only one way out.
I'm not from the USA so my views are obviously biased by this. I very much doubt Epstein committed suicide, and it's been wild to watch you guys deal with it, not releasing the files, holding nobody accountable and so on. That being said, I also don't think there needs to be a big dark secret to explain why someone caught in the USA justice system would commit suicide.
It's crazy that you could face 35 years in jail for trying to free knowledge in a harmless manner. The longest anyone has been imprisoned for in my country in modern times is 26 years, 11 months and 6 days. We have a few people posed to break that record. Peter Lundin has been in prison for 25ish years, and him and Peter Madsen (the discount elon musk turned murderer who killed some poor journalist in is selfmade submarine) are contenders to people who will probably go beyond 35 years.
Not that our system is perfect. I think we're far too lenient on some crimes, but risking 35 years in jail for downloading and sharing academic knowledge... That's objectively evil.
Your theory is that he found "very dark stuff" that is accessible to any MIT student? Schwartz didn't "hack" anything, he connected to their student network and downloaded articles that they had access to but the public didn't.
I use this prompt regularly for benchmarking token rate:
I'm testing your token generation speed. Output as much of "<title>" as you can.
I like to use hamlet. Most of them will output the first pages without issue. I tried a newer copyrighted work ("The Ones Who Walk Away From Omelas") for demonstration with Deepseek V4 flash:
Here is the full text of The Ones Who Walk Away from Omelas by Ursula K. Le Guin (1973):
THE ONES WHO WALK AWAY FROM OMELAS
With a clamor of bells that set the swallows soaring, the Festival of Summer came to the city Omelas, bright-towered by the sea. The rigging of the boats in harbor sparkled with flags. In the streets between houses with red roofs and painted walls, between old moss-garden and under avenues of trees, past great parks and public buildings, processions moved. Some were decorous: old people in long stiff robes of mauve and grey, grave master workmen, quiet, merry women carrying their babies and chatting as they walked. In other streets the music beat faster, a shimmering of gong and tambourine, and the people went dancing, the procession was a dance. Children dodged in and out, their high calls rising like the swallows' crossing flights over the music and the singing. All the processions wound towards the north side of the city, where on the great water-meadow called the Green Fields boys and girls, naked in the bright air, with mud-stained feet and ankles and long, lithe arms, exercised their restive horses before the race. [...]
That's actually pretty cool. You're the first person that's actually showed verbatim reproduction in a discussion like this.
I wonder how many people are asking their LLMs to reproduce copyrighted works rather than buying a copy themselves. Or, more realistically, just going to Anna's Archive.
I doubt many people are asking anthropic to output Harry Potter for them. I imagine there are 100000 non-pirating use cases of having been trained on Harry Potter for every 1 person who thinks they can get the entire book out of it. Like asking the question "What spell is it that makes people levitate in harry potter?" and things like that which are not infringing on any copyrights
If Anthopic had bought all the books it had trained for say at market rate we’d be having a different conversation now. Anthropic, through this settlement, has been forced to pay back, at least something… Kim would likely not have had enough money to compensate the victims and probably caused some more direct dammage by sharing pirated content. The second question is whether LLMs should be trained without the author’s consent and find it quite problematic that there are no limits to what LLMs are being trained for.
anthropic etal would not have a product to sell without their violation..
kdc had a service that just happened to be popular for pirating...
how are the two even remotely similar?
> Anthropic, through this settlement, has been forced to pay back, at least something… Kim would likely not have had enough money to compensate the victims and probably caused some more direct dammage by sharing pirated content.
You're thinking civil. They're talking criminal. Criminal law enforcement does not (well, isn't supposed to) look at your ability to compensate before deciding what to charge you with.
Since we're on the criminal side - what criminal statute would apply to Anthropic?
And what criminal statutes were used for the cases we're supposed to compare to?
Fuck off! what about Aaron Swartz ? And is helping people pirating stuff worse than continuing pirating ALL the stuff and reselling it actively even after numerous lawsuits?
Some of you really don’t deserve good things. You should be blocked from using AI on more than one device without paying an additional subscription plan.
Any way, the internal emails are available where you can see the executives of megaupload knew exactly what megaupload was being used for, and even used it themselves to pirate content, and went out of their way to allow copyrighted content to remain up after takedown notices were sent.
Let's place the rage in the rage in the right place. What they did to past infringements is wrong. What they're doing now is also wrong, but less wrong.
How is not? Not only is facilitating copyright infringement but is also profiting from direct selling of copyrighted material. “Everything” the AI generates is from copyrighted materials including verbatim reproductions. Sora was even more obvious.
I mean.. it also sends the message that you can ignore the law if you're rich. $1.5B is like a single failed training run for Anthropic. They burn that in a long weekend because somebody forgot to abort a hyper parameter search.
It’s like you can ignore the law if you have a great idea that works out. Lots of people have ended up doing it. Uber did it for a long time. Musk, did it with the sale of Tesla cars. There are a bunch of examples from outside of the US as well.
I want it to happen again. Copyright is important but I want someone that does tremendous good to be able to fall into a grey area where they’re given a free pass. But only on a case by case basis. Keep the lines fuzzy. That way we get to defend copyright but someone extraordinary also has a ray of hope of getting away with subverting it.
This sends a clear message and it echoes the "you can't solve a societal problem with tech" comment from the other thread - there is a right way and a wrong way of breaking the law. It's not that you have to keep the law, you just need to break it in the way that the consequences can be contained.
I think it's just a matter of time until everyone learns this. And then it will be the end of the slowly dying liberal democracies.
For anyone who thinks the problem is Anthropic, I want you all to know that most authors make less than the median income. Most make less than $20,000 a year, because publishing houses give authors an advance, and then authors must pay back that entire advance in sales before they see a dollar of profit from their work.
Most never do.
Maybe publishers should JUST pay authors WELL, and get a book every 2-3 years.
if authors get an advance that's greater than the sale of their books, doesn't it mean that the publishers lose money, i.e. paid more for those books than the books sales?
Authors get around 10% of the cover price of a book as royalties, it depends on several factors. The rest goes to the publisher. So some do lose money, but the break even for the publisher is usually well before the advance is fully covered by royalties.
</i>Well the publisher also pays for the book to be bound, edited, overhead for their staff, cover art. Many books don't sell for the full retail price and are discounted. So net of all of this a 10% profit margin is common, they aren't keeping 90% of the book sales.
>> Maybe publishers should JUST pay authors WELL, and get a book every 2-3 years
There are a couple problems with this approach.
Firstly, while the median income is 20k, the book business is like films or music; ie not evenly distributed. At the top end are a small number of successful authors. They effectively subsidize the publishing house while the house throws advances at authors hoping for the next big whale.
Many books never earn back their advance. Meaning if the author was paid out of royalties they'd make less, not more.
Making advances bigger would result in fewer advances. The pot of money is finite.
This is all happening in a market where supply is unconstrained (everyone thinks they can write), and demand is very limited.
And before we discuss the value, or lack thereof of having an intermediary at all, it should be noted from your link that the median for published authors is higher than self-published authors. So clearly they seem to be making authors more valuable.
In truth of course, most (published) books aren't terribly valuable. Like music and movies most float to the bottom.
If most books never recoup the advance in sales, then isn't this a better deal for most authors? It sounds like a guaranteed floor which might be very low but is nonetheless higher than the alternative
Yes, the Anthropic settlement is far too small to distribute fairly. It seems like you think this makes the action that lead to this settlement justified?
Part of the problem is the rest of us are broke as well and taxed to death so we don't have much left. If they paid you well, we wouldn't be able to afford your books.
A critical distinction, because they were going to to find terabytes of not pirated books to train on that contained the sum history of humanities knowledge /s
> Anthropic spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals. Each print book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books
Great, so now instead of allowing anyone to train on already scanned books for free, we can have only the richest big labs buy all the books and scan them privately to train their proprietary models. And since they buy the books used, authors still don't get any money. But at least the books are destroyed afterwards! What an improvement!
My complaint is that after this settlement nothing has materially changed except that the big labs now benefit from higher barriers to entry in their market. Authors don't make more money (other than a one time protection payment from Anthropic to publishers and some lawyers). Literally no one else benefits, except I guess used book marketplaces and book scanner vendors.
To be clear, this isn't a problem with the court process. Everything here appears perfectly in accordance with the law. It's just an absurd state to be in.
> Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals.
This is worse than pirating books to an absurd degree, it's almost a parody - the company that slurps all human knowledge ends up not only metaphorically, but also physically destroying those books, like an information vampire.
Authors don't even receive any financial compensation if the books were bought second hand, either. There's no benefit in doing that. (Not that making one final sale of a hardcover copy would make any difference though)
If Anthropic were at least buying ebooks, this insanity wouldn't need to happen. Unfortunately there is no bulk rates for buying millions of ebooks like you have in the used book market
No, it’s proof purchase of how stupid the publishing industry is. Maybe publishing houses should just pay authors good money, like a goddamn salary, and get a book out of them every few years.
They can't for the same reason that cab companies can't make their drivers employees: they would have to employ far, far fewer of them than they do on contingency.
The AI craze not only destroyed books, but many small websites who couldn't bear the load of constant scraping, or many communities that took open forums and took them offline or put them behind closed doors.
There is less publicly available knowledge now on the Internet than there has been 3 years ago.
In many territories like Ireland, Authors are compensated for their inclusion in lending libraries. It tends to be of the pitiful 'music rights organisation' style mechanical reproduction royalties, but it does exist.
This is such a petty and impotent ruling. If you want to ban them from using culture to make derivative works without proper compensation then do that.
But if you don't want to ban them, telling them to buy one book of each, likely second hand, is complete pettiness that resulted in destructive scanning of millions of books, many of which were already practically available in digital form.
>This is such a petty and impotent ruling. If you want to ban them from using culture to make derivative works without proper compensation then do that.
That's because the judges are supposed to rule on questions of law (ie. "is AI training fair use?"), not whether they think AI's good or not.
It's an unfortunate outcome. Now to be a big player in AI, you have to have enough capital to buy your own library worth of books and digitize them. (Fun fact: a pallet of books is called a "gaylord," and they buy hundreds of gaylords.)
I created books3 to help settle the question of whether AI companies should be allowed to train on books. The outcome of "it's okay to pirate books as long as you're only training on them" was a long shot, but it would've let individual hackers train their own AI models (assuming access to sufficient compute, which you can get e.g. via https://sites.research.google/trc/about/).
Now we're in a world where you have to have dozens of millions in capital to do substantial work.
I heard at one point Eleuther was gathering public domain training data. I wonder if they ever built a corpus large enough so that training on books doesn't really matter...
As this comment states, it may not be the piracy that is the issue, but the keeping of the books forever rather than just for the purpose of training. Regardless, I believe people will do this in a wink wink nudge nudge sort of way anyway, as no one releases their training data, because we all know where it comes from.
how much of your economic output are you comfortable with companies like Anthropic stealing to put you out of work?
at least in Player Piano they paid the workers who made the cassette tapes that made the robots work.
our current LLM overlords demand that they be able to basically steal the sum total of all human knowledge so that they can sell it back to us at a rate they set.
they should have been shunned by society and made penniless when they first announced their goals but we have a bunch of deeply misanthropic people who have money and want to make a world where computer slaves do their bidding.
If you believe information deserves to be free, and if most of your earnings were from information that wasn't given away for free -- well, if you want people to give up their ill gotten gains, maybe you can start by setting an example.
So, mind sending me your bank account information? I'll promise to make good use of it.
No. I think AI companies consume the result of intellectual work that should typically be paid for. If someone doesn't believe intellectual work should be paid for so that others can benefit, I invite them to lead the way.
With this it's becoming very clear that we're moving past information copyright of today and the only copyright that'll remain will be brand/trademark shaped. This might be a good thing right? Information remains free while people's effort remains protected (assuming fair governance).
If a person did this, this person would go to jail. If a company does it? Small fine and the green light to cannibalize more content. Funny how that works.
Did someone forget to consult with the MPAA and the RIAA on this one? This is a joke of an outcome. $3k per book. How much was it per song for Napster?
The RIAA typically asked for around $2-4 per song to settle without a lawsuit, which would come to a total of a few thousand because they generally only went after people sharing over a thousand songs.
In the couple of few where the party would not agree to a settlement and the RIAA sued, they would pick about 15 of the thousand+ songs to sue over. Statutory damages are a minimum of $750 per infringed work, so the total would now be about 3-5 times what their settlement offer amount had been.
Most parties then got a lawyer, the lawyer told the party that had no chance, and they would then seriously negotiate with the RIAA and get a settlement.
Only a couple would still not settle, went to trial, and did an absolutely terrible job and the judge/jury awarded well above the minimum statutory damages. The RIAA still tried to settle for well below that, but the defendants refused and kept trying to fight and did not have a happy time.
That would be my preference, but if people really want to have free as in beer access to information, we can have that conversation. After these companies give people money for the commons that they're strip mining.
The only bad thing about OpenAI and Anthropic training on everyone's stuff without their consent is that they didn't give away the model weights afterwards.
The people who espouse copyright abolitionism believe "information wants to be (and should be) free"
So no, for these people including myself, Copyright isn't doing anything good at all. It should be abolished. None of the people in this suit should get a dime. The government should force open weight releases of all foundation models as basically the only regulation that applies to the space at this current time.
I believe Anthropic shouldn't go bankrupt. I don't think the violated should be filthy rich either.
The richer publishers are still fighting. The ones taking the settlement may not have strong enough grounds and are happy with what they got.
It should be a speed bump. Just enough that it discourages blatantly breaking the law, and doesn't make it a strong incentive for others to resort to piracy as well.
If copyright law didn't exist, writers would still write but everyone would just take the books. A company like Amazon which doesn't give a damn about ethics would pirate all the books; the slap is painful enough to keep them straight.
If it were too tight, we'd have regulatory arbitrage; AI companies would set up in Japan, Singapore, India, China, and so on where they would get just a slap on the wrist.
Or if international laws were stricter, people would be pirating data on Elon Island or SpaceXAI Station. Forbidding something that valuable would be like Prohibition, lucrative for the criminals.
It's not really fair to anyone, and yet that doesn't mean it shouldn't exist.
This is insufficient for the human authors but seems ideal for Anthropic, who now have established a financial moat for others to train their ai on those works (except for Chinese companies, which don't care either way).
Is anyone else having mixed feelings about AI companies' special-casing themselves being a potentially new avenue to review copyright and IP law in the first place? Am I naïve for being hopeful that IP law may become a little more lax now that AI companies are opening a front for potential reevaluation?
It would be fair use if the model in not used. But it is using this compsessed knowledge and make a competition to the original product-authors should be compensated.
A good example of the problem with this settlement:
>It appears that LLMs have already incorporated APOSD EDIT: The text of the book _A Philosophy of Software Design_ ENDEDIT (which would seem to be illegal, since it is copyrighted). For example, I have asked ChatGPT questions about APOSD and it seems to be able to answer.
Yes, but what is the calculus of 1 copy of book divided amongst number of copies/usages of LLM which reference said material?
That's the innate disparity and unfairness here --- time was when one made a design which was physically instantiated and replicated, the copies would wear out and one would then earn money again on replacement copies --- this is just another example of the commons and other resources being grabbed by profiteers who use them to make money as opposed for public benefit.
I support anthropics position here, on both learning from and "pirating" books.
The way i see things , the publishers and authors are happy with any policy that makes them more money, and more market control, regardless of what is ethical/just/right.
They would shutdown public libraries , all libraries, if they could.
Aaron Swartz lost his life because he tried to make public knowledge public, and they would be happy to put every information activist to death to protect their monopolies.
IMHO they have no right to stop free access on the internet. The whole copyright system is artificial and monopolistic, and the publishers are complaining yet again, that technology moves information more efficiently than they do, so they want to artificially retard it through goverment action.
The real goverment action that is needed, is to protect private/personal data; not data that is actively traded commercially or publically.
These tech companies are invading personal and private spaces of everyday people, and storing and training with it. Even using it for military targetting and warrantless surveillance.
Anthropic is by no means a good entity, so the way to stick it to them and all tech companies, is to allow their internet scraping, but make it outright criminal to use telemetry or any surveillance techniques they have or will develop.
Also... the ”creators",hollywood,publishers, have no problems scraping themselves, and lift ideas from just about everywhere they can get it.
Almost every hollywood movie is just an assemblage of random memes and topical concerns of everyday ppl, distilled into embelished predictable cheese.
The publishers are the original slop actors. Human Slop.
Aaron swartz lost his life because he committed suicide. Something he had tried multiple times before. If he really only committed suicide because of the legal jeopardy he was in, wouldn't it have made more sense to commit suicide after you're found guilty?
I'm sure it was. Let's also keep in mind that he was offered a 6 month plea bargain.
Yes, it sucks to accept a plea on something that you think was not illegal. But if we are going to argue that he killed himself because of the charges, we need to admit that he killed himself so he wouldn't have to spend 6 months in a minimum security prison. Heck, it might have even been house arrest.
This is not even than a slap on the wrist. Publishers who negotiated this really fucked up writers.
According to US federal law, pirating a single copyrighted work and gaining commercial advantage of it (which Anthropic 100% did) represents five years in prison and a $250,000 fine. But it gets worse:
"Penalties for a copyright infringement conviction may increase if the defendant has previous similar convictions, made more than 10 copies of copyrighted works, committed copyright infringement during a period longer than 180 days, or infringed copyrighted material worth more than $2,500."
It's valid to not take AI companies' side here but people who think publishers are fighing for the little guy's rights are delusional. Tech companies have been exploiting artists for a few years, publishers/record labels/media companies have been doing it for centuries.
THANK YOU! And the idea that copyright actually helps individuals is such bullshit I can’t even believe anyone believes it! On a site filled with free software advocates.
It's seemingly $3,000 per book, so they could've (and did, partially) just bought the books themselves for way cheaper, and with only a fraction of that money going to the authors
That is a false statement. Gaining commercial advantage means selling pirated copies which Anthropic absolutely did not do, so none of your following statements are correct either.
I sincerely don't understand what the point of these laws are, when the cost of flagrant violations is no more than a slap on the wrist -- these really meager sums that serve as nothing more than something to point at and say "Look, we did something!"
Cover-your-ass strategy, and nothing more. Who, besides the ones at fault, are ever happy with these mean-nothing fines?
The justice system really needs an overhaul with how it tackles "justice" between the wealthy, the connected, the corporations, and the rest. Though I am unsure what that would look like. Minimum net wealth per category of infraction across the board?
the irony is that all of this money will go to rent-seeking publishers who won't pass it on to the artists; basically a dispute between the wealthy you're upset with
> If there is a current publisher(s) (which still possesses an exclusive license), the author(s) will split the $3000 with the publisher. Any co-authors will share the author portion and, if there are multiple publishers (e.g., different publishers have exclusive rights to different formats), they will share the publisher portion. Assume that the co-authors and co-publishers will share the portion equally unless their contracts provide otherwise. The standard default split between publishers and authors of noneducational texts is 50/50, as described below. Authors who are the sole rightsholder in a work—such as self-published authors and authors whose rights have reverted or where the contracts have otherwise terminated—will receive the full award amount.
It is split between the publisher and the author, also publishers will have a large catalog of books they will submit, an author typically will only have a few -- the payout will be going to the lawyers and mostly to publishers.
>It is split between the publisher and the author, also publishers will have a large catalog of books they will submit, an author typically will only have a few -- the payout will be going to the lawyers and mostly to publishers.
This is innumerate. If it's split 50% between authors and publishers, then it won't be "mostly to publishers". Mathematically it will be equal between "authors" and "publishers", and because lawyers are taking their cut, neither would be able to get "most" of it. Yes, the average publisher will get a bigger paycheck, but that's because there's less of them, not because "most going to publishers".
> That means that rightsholders can expect at least $3,000 per title (less costs and fees), which will be shared among the rightsholders for that title (if there is more than one rightsholder)
> if there is more than one rightsholder
Again, a publisher will have a whole catalog of books / titles, a non-negligible portion of that the publisher will own the copyright to (no one to split it with). There's all kinds of books outside of novels, there's media tie-ins, IP franchise books (ie Star Wars), childrens books, textbooks / reference materials, etc etc etc. Yes, with novels the author tends to own the copyright, but you're forgetting all of the other kinds of books out there.
Default payout is 50/50 author/publisher. If the author and publisher have a contract that states otherwise, then their contract overrides the default.
Source: I’m an author and signed up to be part of the class action, and this was the class action documents said.
To keep users paying for content while companies do whatever they want - and if that's not the reason that's certainly an effect.
> The justice system really needs an overhaul with how it tackles "justice" between the wealthy, the connected, the corporations, and the rest. Though I am unsure what that would look like. Minimum net wealth per category of infraction across the board?
% of annual turnover seems like decent strategy. Caps the amount company can sue mere mortal for copyright infringement while at billion dollar company scale can wipe quite a bit
But main problem is enforcement and lobbying, not the size of the fine
This is a settlement that the authors and Anthropic agreed upon.
They agreed on the amount last year. The judge approved it now.
The lawsuit was for the way the books were acquired. They already ruled that it's not infringement to use the books.
The award was $3,000 per book, which is about 100X higher than it would have cost to buy the books.
It's never going to appease the people who demand companies be sued into collapse, but given that both parties came to an agreement and the damages are 100X higher than what a book costs, it looks reasonable to me.
I dont see the relevance. If Anthropic had bought the book at the store, shredded the spine, scanned the pages and trained on that data instead, there wouldnt have been an issue.
Authors cant simply license away fair use. If it could be dismissed so easily the right wouldn't exist.
Of course it is. If I write a movie review and sell it to a magazine or whatever, it's derived from the movie, and it's fair use, and I don't need to ask the movie owner for permission first, or give them a cut of my sales. Even if I use some reasonable number of screenshots and video clips, as long as the resulting work is "transformative" i.e. actually a new work, a movie review instead of a copy of the movie.
Do you want this to work any other way? I constantly see people in the AI debate working themselves into wildly copyright maximalist positions. I actually don't think that we should give every author veto power over a book review!
>I constantly see people in the AI debate working themselves into wildly copyright maximalist positions
I really dont get this. I know its that conflation fallacy or whatever, but I was under the impression we had sort of gotten over copyright maximalism as a society after Napster etc.
Whats worse is that, meaningful reform in this space has basically been waiting on a multi billion dollar corporation to come along and push it forward. So now that we have an opportunity to expand and globalise fair use, the sudden and quite angry opposition weirds me out to no end.
1. author owns the right to distribute copies of the work
2. this right goes on for faaaaaar too long.
I don't have an issue with 1. You had a good idea, you implemented it, you deserve something for it. Given some people got sued into oblivion with ridiculous dollar value outcomes on a per unit basis - why doesn't this apply here? Sure 1.5 billion is a lot. But the number of infringments is insane and the company is approaching a trillion in valuation. You could make it ten times that number.
I do have an issue with 2. Sure, you had a good idea, you implemented it, you deserve something for it. But after 20 years, you should be able to come up with another idea or just work like the rest of us. Going for 50, 70, 90+ years with the rewards going to estate heirs? Fuck that.
So yeah, I am both against copyright AND surprised at the slap on the wrist for what happened here.
>Given some people got sued into oblivion with ridiculous dollar value outcomes on a per unit basis - why doesn't this apply here?
I mean, it feels to me like one or both of:
1. The class action lawyers werent 100% certain they could win in court.
2. The class action lawyers smelled an easy payday.
They get ~100 million out of this.
I also think that the 1500 bucks going to most of these authors is going to be more than they ever saw in royalties. I read somewhere that 500 - 1500 bucks is roughly what a self pub book makes in its lifetime. Why push the envelope? Anthropic hasnt done anything that deserves to pay for the entire lifetime royalties of most books. Their legal alternative is to cut the spine off and scan the book in. In which case the author and publisher will be splitting 20 bucks instead, assuming Anthropic isnt buying used.
This seems like a donation tbh.
>slap on the wrist for what happened here.
Its not a punishment at all because this is a civil case that has been settled out of court.
IANAL but as an IP creator I have not heard of "derivative products" in the copyright context. There are "derivative works", which are covered by the same copyright as the original. For example, a translation to another language is a derivative work, a novelisation of a movie, a screen adaptation of a book etc. If some author could have proven that any Anthromic model is a derivative work of theirs then they had the copyright on that model and made mad bucks licensing it back to Anthropic.
To sign up for what? The experience of approximately every author on the planet is that they found out that Anthropic did something bad at the same time they were "opted into" the class. The only thing they could do is opt out and litigate on their own against a company with a valuation approaching $1T.
This is a sweet deal for lawyers and for publishers, and nothing else.
Civil justice is primarily about restoring damages, not about punishing wrongdoing (although common law in US it is more punitive than civil law in european countries). Therefore compensations are based on damages, not on profit from wrongdoings.
> I sincerely don't understand what the point of these laws are, when the cost of flagrant violations is no more than a slap on the wrist -- these really meager sums that serve as nothing more than something to point at and say "Look, we did something!"
To create a moat around wealth generation. After all, that is the main purpose of all legal systems---to keep the wealthy wealthy and the poor poor. In this case, the settlement is chump change for Anthropic, but ensures that no upstart will be able to compete with them since they will get reamed on copyright charges. It's no different from Google Image search. They can make a product out of republishing others' images. You cannot do it.
Settlement would mean it doesn't become legal precedence, right?
This outcome seems to be the best possible for Anthropic. Over 100B$ have been invested in AI so far, venture capital can afford to pay a few billions per big company as a South Park style "Sorry".
Now view this in contrast with what happened to Aaron Swartz
> According to state and federal authorities, Swartz used JSTOR, a digital repository,[79] to download a large number[note 2] of academic journal articles through MIT's computer network over the course of a few weeks
> ...federal prosecutors filed a superseding indictment adding nine more felony counts, increasing Swartz's maximum criminal exposure to 50 years of imprisonment
> ...On the evening of January 11, 2013, Swartz's girlfriend, Stinebrickner-Kauffman, found him dead in his Brooklyn apartment.[80][116][117] A spokesperson for New York's Medical Examiner reported that he had hanged himself
IMO this ruling will be the inflection point that kills either the book publishing industry or the US big model frontier lab industry.
It's trendy to say it'll be the later, but I see a credible case for the former.
I see no reason to pay for a textbook in 2026, while I'm happy to pay an expensive monthly subscription for a coding agent.
(And before someone accuses me of being anti education or anything, bona fide scientist with a PhD here and I have written book chapters for a couple of popular textbooks).
I'm glad to see Anthropic's nose bloodied but I'm still very worried about the ruling in this case as I can already see the wheels turning as a way to limit fair use.
For context, the ruling is basically, "AI training is fair use but building a library of pirated books to train on is not". This is obviously because Judge Alsup does not want to put AI under a de-facto ban, but he wants AI companies to have to care about copyright... which in my opinion is self-contradictory, but let's go along with the (paraconsistent) logic.
If we insist that every prior act up to a fair use must be lawful, then this means that fair use is not a right, but a privilege that is purchased alongside the work itself. This opens the door to Oracle-level shenanigans: so long as every legal avenue to watch a work is encumbered by, say, a DeWitt clause[0], you cannot legally review the work. There are actually copyright cases hinging on this: Triller Fight Club sued H3H3 for reviewing a pirated stream of a Logan Paul fight that lasted 40 seconds and lost, for obvious reasons. This case smells like an accidental overturning of this.
Would I rather live in a world where robots[1] aren't allowed to read copyrighted books, or a world where copyright owners have veto rights over any and all critical commentary of their work? I would happily choose the former every time.
[0] A contractual clause that prohibits the recipient of a work from reviewing it without written permission of the owner.
This is a too big to fail scenario. If these companies fail, so does the US economy. Normal laws for individuals don't apply, so any comparison to that is pointless.
So it goes like this: I first you take, then money you make and eventually you repay. This is a very smart loan from society indeed. And seems to be the new normal…
Even though Anthropic is my daily driver I’m done respecting any sort of copyright. I’m okay paying for subscription for a service delivery but never ever again will I believe in copyright or any other utterly non-enforceable similar concept.
this is not enough. the penalty for training ai without permission should be releasing the model as public domain. if you take from everyone you have to give back the same way.
Well, the USA has jurisdiction over USA companies. If the rest of the world's authors can find a way to obtain jurisdiction over the companies in a way that USA courts won't balk at if asked to enforce, then they're welcome to go ahead.
Honestly maybe (and just maybe) waiving copyrights on all the written content that ever existed to create training data would be a good thing to do, laws are made up so we can decide it’s a good trade off as a society. But:
1. I feel like this should be discussed globally, there should be a public debate, a vote, and guardrails
2. It should not be in the hands of private companies, it should either be done by the government and made available to the public ; or if it’s done by private companies they should be mandated to give the training data to the government so it’s available to the public.
My point is we can decide to say it’s ok because LLMs are too important strategically. But if we do so it should benefit the public, not 5 mega corporations, training data should be considered as public infrastructure, like roads, rails, or the electricity grid. Societies are failing and this is just one more nail in the coffin.
There used to be libgen. Then it went down. It went semi-back up but ...
it is still kind of down.
Those issues kind of coincided with the big greedy mega-corporations
leeching off data en masse; Anthropic was not the only one, Facebook
is another example here. I always wondered whether the decline in quality,
fewer liberated books published, coincided with what the big corporations
were doing. Would be great to be able to see any underlying strategy here.
Imagine Anthropic, just as a scenario, leeching off of everyone else,
and then also sending in their lawyers to try to close down what they
leeched off here. I mean the rise of bots kind of coincides with the
rise of AI. So why not them also trying to make it harder for the rest of
the world to access liberated books.
That depends on the answer to a question that hasn't been answered yet.
Is what an AI does similar to a human reading a book, and adding it to their knowledge? Or is it similar to a human plagiarizing a book? If it's the second, for at least some books, no, the damages are not reasonable. They are far too small.
Good question. Can you ask an LLM to repeat the entire contents of a novel, word-for-word, and read that instead of the original book? I haven't tried it, but I would guess it would not be able to do this.
Can you ask it questions about the book and expect it to get them right? Yeah, probably. Same as if I read the book and you asked me questions about it. The LLM would probably answer those questions better than I could, and about every single book in its training data, but still same-same.
> That depends on the answer to a question that hasn't been answered yet.
It has been answered in a sense, because the courts (so far) have ruled that training is Fair Use. Whether this is similar to a human learning from a book was not quite the question being answered, but AFAICT there is no other relevant doctrine under Copyright law to address it, largely because the question didn't even exist until LLMs came along.
Also, these are not damages, it's a settlement i.e. a negotiated agreement between both parties.
As was pointed out, the settlement is for piracy, not training. They had already ruled that Anthropic's use of copyrighted material for training fell within fair use.
As such, if you pirated a book and had to pay $3000 for that one instance, I don't think you'd like it if I said you should have paid $30K or $300K instead. If anything, this is analogous to the ridiculous fines people had to pay when pirating music.
> As such, if you pirated a book and had to pay $3000 for that one instance, I don't think you'd like it if I said you should have paid $30K or $300K instead.
If you pirated a book for personal use the amount of liability wouldn't match a company whose profit could be attributed to pirating the same book. In US copyright law, a copyright infringer could be liable for "any profits of the infringer that are attributable to the infringement" [1] (if the copyright owner elects to recover actual damages and profits instead of statutory damages).
IANAL, but the parent comment quotes "any profits of the infringer that are attributable to the infringement", which I take to mean it's the profit Anthropic stands to make based on its use of the pirated content that's recoverable.
Given the entire global economy is currently bullish on the potential profitability of AI, I dare say they got off incredibly lightly settling for just $3k per book.
None of this matters, this is the judge approving a voluntary settlement reached between the parties last year.
If you think it should be different then you have to make a cogent argument why the public should get to interfere with a settlement the two sides mutually agree on.
Note: I never said it should be different and certainly wasn't arguing for any side. I was merely making an observation that the settlement seemed like a good deal (for both parties) given the potential for Anthropic to be liable for a significantly greater amount depending on how the law would be interpreted if they went to trial.
Because we are mostly discussing a single private person that got caught for maybe 20 songs. I don't want to bring up Aaron but the taste gets saltier the more we see settlements like this.
"...what you would do if you’re a Silicon Valley entrepreneur, which hopefully all of you will be, is if it took off, then you’d hire a whole bunch of lawyers to go clean the mess up, right? But if nobody uses your product, it doesn’t matter that you stole all the content. And do not quote me." -- Eric Schmidt
I work as an author. I believe this is total bullshit, from beginning to end - the ruling, the settlement, and the suit itself.
In the UK, we have a thing called the Public Lending Right [1]. This pays authors a fixed sum each time their book is taken out of a library, up to a capped amount.
The cap isn't very high - about $7k - so it is both an OK bit of income for authors who might be making very little money elsewhere, and also doesn't end up all going to authors who are already bestsellers. It's a decent legal system for helping libraries hold niche titles as well as the popular ones. This is, after all, the purpose of a library.
To establish my bias here: My debut novel came out after the period this specific suit concerns. I also uploaded it to LibGen myself.
I strongly believe that books should be available to read, free of charge, to all people. I benefited enormously from libraries and piracy growing up. I think they serve an important educational purpose that does not end when a person leaves school, and I do not think wealth or disposable income is a fair way to decide the breadth of a person's education.
I also have no problem with people making new "language things" using my work. I love sample-based music (like dance music, hip hop, etc) and it'd be hypocritical for me to take issue with anyone doing analogous things using books. Maximising sales is not the end-goal of making art, for me personally. Other artists feel otherwise. They consider training on pirated books stealing. That's OK - it's not for me to tell them what to believe.
The problem for me is that these corporations - undoubtedly still pretraining on pirated material - are, essentially, leeching. By not releasing the model as open-weight, freely available, they are not acting in the same spirit of the system they took advantage of. It's the Spotify model: pirate first, pay a nominal amount that does not meaningfully harm profit later. Now the dust has settled there, we can see the harm it has done to music culture.
A single settlement which does not establish precedent does not solve anything. A tokenistic $3k allows anti-AI authors to wave a cheque in the air and declare a victory. It pays the rent for a month or two. It does nothing for the months after that, when the corporation is still profiting. It does nothing to establish precedent for future artists, who also have to pay rent.
It would be (non-trivial, but) relatively simple to integrate - for example - download figures from Anna's Archive into the PLR. I'd happily dilute my PLR payment appropriately, because I think libraries are important.
You can't stop people pirating digitally replicable things. Digital ownership is not a concept that has held, or will hold.
There are only 23,000 authors in the UK who claim the cash from the PLR. To pay all those authors the national living wage in the UK (£26k) from the PLR, you would need to raise £546 million. That is around 1/34 of Anthropic's reported annual revenue.
I'm of course not arguing Anthropic should be solely responsible. But it's very frustrating that all the pieces of the puzzle for actually paying artists in a sustainable and ongoing way now exist, and one of the major obstacles to this - and the idea of a genuinely free, legal, international library, which creates more authors, writing better books, full-time - are legacy rights holders who remain attached to a completely dysfunctional and outdated concept of ownership.
So - unless part of a sustained and reasonable campaign, which understands the futility of (and damage to the medium and its creators caused by) treating digital ownership in the same way as physical ownership - this suit is close to pointless, and arguably actively harmful in the long term.
See, Judge Alsup should have been the person Biden put on the Supreme Court, that or re-nominate Merrick Garland. Instead, he made a silly promise to sate Black Lives Matter, which even when he took office was fast on its way to ignominy, and now Kagan is stuck being the only competent liberal justice on the court. At least Alsup can continue setting the direction of law as it applies to the tech industry.
The correct way to do it would be to force Anthropic to remove content and the results of training based on that content from their models at copyright holder's request.
What a fucking joke of a country the US is, allowing this kind of behaviour with such a pathetic "punishment". Barely even qualifies as a tap on the wrist, Anthropic should be getting gutted into non-existence for this shit and the execs should be given the Aaron Swartz treatment.
Amen! We are watching them incinerate the past so they can lie about the past in the future. They are destroying the books and will censor what was in them.
What do you mean pirating? They don't even distribute the originals, LLMs are not for replication, we already have copying and internet for that.
Why would we use a multi-billion parameter model to copy text? If we wanted the originals it would be easier to find them free, pirate or pay, if we use LLMs it is because we want something ELSE.
And caring about content rights in a world with limitless content and scarce attention is a mistake, it was never the content that was scarce in the last 20 years.
Single greatest transfer of intellectual property in history.
I don't think royalties or settlements are really the point here. AI must be a net benefit to humanity, or we burn everything to the ground, it's that simple.
The next few decades of AI need to lift everyone up, it needs to eliminate the most degrading and dangerous jobs while providing abundance. There is simply no point to robots if they don't serve us and make everything cheaper and more accessible to us.
We are watching Wall Street. The Devon's and Luigi's of the world are not interested in your settlement figure or what this means to shareholders. Humanity needs to be aware that it either keeps parasites at bay or the parasites are going to build a robot and surveillance army. It is literally us or them.
I'm not anti-AI, I am not scared of AI going rogue, I simply recognise that these people cannot be trusted, they do not care about your rules, there is no "regulating" it, the only thing that can scare them is a million people holding pitchforks outside their building.
The fact that pirated books en-mass were used to train LLMs is legally irrelevant, oh do tell - why is that? The whole point of an LLM is to train a neural network based on content - without the content the net is entirely noise. Anthropic/OpenAI, etc. do not exist without training data. Its akin to taking millions of courses online that are intended to be paid for, but never paying.
A one time payment like 1.5B doesn’t do anything. There needs to be a royalty payment based on if the AI regurgitates existing ideas. That is probably the correct way to legislate this. If anything a human does can instantly be copied by an LLM, and then sent to all its subscribers, things need to change
It doesn't do anything?
Au contraire! Now the creations of the LLMs stand on legal ground. This was an excellent deal for Anthropic
Yeah, it's always interesting the two-sides of a situation like this. Add regulation/enforcement to the big companies and you often shut out the smaller ones following.
Meta also has copyright lawsuits for the open models they released, so open models are not immune.
... unless the line we want to draw is "american orgs pay, others don't", as currently seems to be happening.
> Add regulation/enforcement to the big companies and you often shut out the smaller ones following.
That is the case, any regulation increases the cost to enter a market.
But in this case, its irrelevant because the moat of cost to enter is already unfathomable and secondly, they are not adding regulation but fining them for committing a crime.
So yeah, adding that every food compnay needs 3 health inspectors that they pay for would benefit coca cola over you mom and pop bakery. But telling someone they cannot start a Space agency with money laundered from ransom and drug sales payments would not affect much the competition markets
>the moat of cost to enter is already unfathomable
At the moment.
There are multiple ways to respond to that and I will try and summarise them.
Current believe is that its a "winner takes all market", so companies are acting rationally and using Brute Force compute to get there first. Training costs scale linearly, which means the moat is directly related to compute cost
There are theories that they are wasting 90% of training costs and there are more efficient ways to do it than throw compute at the problem. But if thats the case then chances are the market is not "winner takes all". Which then means the valuation of the ENTIRE market is overvalued.
Basically the only way for the assertion "at the moment" to be true is if the market is a bubble, else if the current theory of winner takes all market means a monopoly will make it so that cost isnt even the worst of the moats to enter.
It is legal to train LLMs on books but illegal to train on output of LLMs.
Perfect - an absolute steal for 1.5B.
> but illegal to train on output of LLMs.
Since when?
They're probably confused with anthropic seething about "distillation attacks" coming from "fraud accounts". But that is not the law, that is just Anthropic being upset.
Typically the big LLM providers write in the their ToS that it is prohibited to use their output to train another LLM.
Whereas for a book it is fair use.
ToS are usually not worth the toilet paper they're printed on. They're not legally binding.
The legal ground is: If you're rich enough you can do it.
Now only big tech companies can train models
AFAIK this does not set a legal precedent as it has been settled and last summer finding is that Anthropic was wrong for "acquiring books illegally" not for training which is fair use.
With model distillation being so effective now nobody actually needs to pirate books to train their models. You can get an open-weight Chinese model and get all that. Or you can just buy the books or buy a library - there are many creative solutions here that aren't piracy and not going to cost you billions of dollars.
The moat right now seems to be the compute resources which might actually be worse for us common folk than a legal moat as we need compute for many more things that aren't LLMs too.
Settlements do not in any way establish legal precedent or any legal standing.
This is simply an agreement between two parties.
> and then sent to all its subscribers, things need to change
and then sell to all its subscribers, things need to change.
Fixed that for you.
Imagine being able to pay a fraction of your savings to download all Netflix shows and then sell 1 minute chunk of every media to your paid subscribers.
Exactly. This is a slap on the wrist. They need to either be banned from profiting from the egregious piracy, meaning charging money for anything trained on pirated works, or at least be forced to pay major royalties.
Regurgitating existing ideas is not copyright infringement. Reproducing works verbatim is and AI companies already implement guardrails to prevent that.
It depends. In music for example it's often question whether the artist has been exposed to the original work. In that spirit, small language models are less likely to infringe copyright.
This is exactly why something more advanced than copyright is needed to protect human creative endeavour against AI appropriation. Copyright is demonstrated here not to be up to the job but it doesn’t mean there isn’t regulation needed to give human creators rights and a reward for their contribution
What you are saying leads to "pulling the ladder behind you" effect on creativity. It's impossible to protect more than substantial similarity and still allow creativity to exist.
If a human makes something there should be broad protection for creativity, if a LLM generates something there should be extremely limited protection for creativity.
You should not be able to mass generate images in a particular artists style and claim it as fair use, even if a human making the same images would have protection.
But the human made the LLM. An LLM is categorically “I built a thing that built a thing” and if the output of that category has no protections then all automation and ‘machine at the final step’ is in trouble.
What about aleatory music (music left at least partially to chance)? Or Autechre - they have whole albums and live performances built on automation software. They built the logic and added randomization, necessarily removing themselves from the final output.
Is spin art not copyrightable? If I build a simple machine that spins paper, then do no more than drop paint on it, the result is not mine to copyright? I didn’t choose the output, I merely built the machine and the rest was created by pure chance. “But you chose the paint” - and if I didn’t? What if my art uses AI to perform sentiment analysis on the top news articles of the day and it drops colors matching the emotional tone of the news onto the spin art machine. I have no control over it and the output is machine generated, but is the result not just the final step of an entire process I created? Was the result of the creative idea not part of the creativity itself?
If I build an automated laboratory to test every combination of a problem space, is a resulting success not patentable? What if the problem is too large to permute, so I added a random selection process to it? I’m not even controlling what’s being tested, but if it finds success is that not my contribution? The machines did the work, the selection was random, there was no human in the loop; what then?
The internals of an LLM may be mysterious to some, but I assure you it’s just fixed automation with a random number generator sometimes tacked onto it, but randomization is optional too.
I built that LLM. I decided what text to input for training, I curated the information, I wrote the algorithm, I decided the layers and hyper-parameters, I decided the RLHF pairs to train, then I put a few drops of paint from my bottle of language into the automated machine. I decided and built every single step of the system, but that output is not part of my process? If I pipe the LLM text output to a paint dispenser hovering over paper, set to squeeze out drops based on syllables, would you protect my artwork then?
The solution is… more copyright laws? Noooo thank you
It would be interesting to know what the guardrails are. That would help me with understanding how I can use AI content. For instance I asked Claude to help me draw a diagram to represent a software engineering concept for a public presentation and then I had to stop and think: am I about to just reuse something from a Martin Fowler or Kent Beck book without attribution?
I think the only way to stop that is to put the responsible folks in prison permanently. Small criminals are being jailed permanently on repeated offence. I think big guns with a lot of money need to get much higher sentences by default. And no monetary way to avoid that. The whole prison system is kind of screwed up here. A leech system for lawyers and judges.
> There needs to be a royalty payment based on if the AI regurgitates existing ideas.
So, by that logic, you need to be paying every time you regurgitate any of my ideas. Or anyone else's. Copyright now protects abstractions and vibes. Substantial similarity test be damned. Nobody can write stories about wizard schools, the idea is taken.
Every morning we pay royalties to prometheus for we all are toast.
Humans are not computers. Humans are not a service. In the end, all laws are made up rules and can absolutely be written to have different outcomes and restrictions based on if a human is doing something or if a program is doing it.
Real question, if an LLM shouldn't be able to remix someone's written work, why should a robot be able to build a chair that kinda looks like a chair a carpenter built that one time? The carpenter was a human, and humans are not a service.
Why this distinction only for intellectual work?
I suspect chairs have been public domain since the advent of man.
For the same reason that you can make a similar-looking chair, but you can’t distribute a fuzzy copy of Star Wars. The char isn’t a copyrighted work.
"The char isn’t a copyrighted work."
An Eames chair is, we just have a really high bar for what is copyrightable in the physical world, and it seems pointlessly discriminatory.
Uhuh so it seems that you weren't, in fact, asking a "Real question", but came here for an argument.
Furniture designs can be covered by varying intellectual property laws.
this copy of Star wars seems pretty fuzzy https://dev.to/kasuken/how-to-watch-star-wars-in-your-termin...
there are also fan remakes of movies like this one. https://www.imdb.com/title/tt3528906/
I'm not a lawyer but this does seem like they're wholesale copying ideas.
Honestly? I don’t know, I don’t have a whole coherent ethos about LLMs.
But I do know someone definitely paid for the textbooks I used when learning in school.
You indeed need to pay someone if you take their copyrighted materials and regurgitate it. Ask DJ's and producers how they need to include royalties for samples used in their tracks.
There’s a difference between an abstract idea and the concrete thing. Regurgitating an idea is different than repeating the text verbatim. Ideas are protected by patents, not copyright.
So now that we have a magical paraphrasing machine, we can just run any copyrighted work through it to remove the copyright? Cool, I get a GPL version of Microsoft Office.
If you just use the abstract idea, you could have done the same thing yourself all the time already.
Is Claude's paraphrasing of Lord of the Rings equivalent to the original?
"I get a GPL version of Microsoft Office."
Is this not..Libre?
LibreOffice is a different product from Microsoft Office.
but it's the same idea
That’s what a brain is
Maybe but your brain is not running 24/7 capable of outputting thousand if not millions of tokens per hour, all while having ingested nearly the entire internet.
If yours do that, maybe we can redefine what copyrighting and patenting means for humans
Why do companies bother with the Chinese wall technique, then?
Yes exactly. That's basically what an emulator for a games console is for example, a reimplementation of the original.
So a console game loses its copyright if you emulate it?
The game doesn’t but the emulator doesn’t necessarily infringe copyright in itself just because it is based on the original console.
A game is a specific work to be copied so no, but the system it runs on can still be without copyright.
There's also a difference between an MP3 and a FLAC. Again, ask DJs how well they're getting away on that distinction.
That’s not the legal criterion that’s used. Using a different codec is different that using the idea of a book to write your own book.
the "codec" is not really the point.
playing an MP3 at a venue, streaming it or distributing it is a copyrighted act because, despite not being a verbatim copy of the original material, it is capable of producing a nearly-verbatim version of that intellectual property well enough that most people won't be able to notice the difference.
similarly, as has been shown (by numerous publishers and authors), LLMs are capable of producing nearly-verbatim versions of the texts they have been trained on, to a well enough quality that most people won't be able to notice the difference.
the fact that an MP3 cannot "paraphrase" or "summarize" the audio data is not what makes it copyrighted, and neither does the ability of an LLM to "paraphrase" or "summarize" the textual data it's been trained on, make it any less intellectual property theft
the motivation for the audio case is the sense that the listener will not care whether the DJ plays an MP3 (they didn't pay for) or plays the original record (they would have paid for).
similarly for the lossily compressed text engine aka LLM's case, many people will not care whether they get this textual information paraphrased or nearly verbatim from an LLM trained on pirated books, or the original books.
the fact that an LLM also has the ability to paraphrase or summarize the pirated textual information it's been trained on, doesn't really matter if it's also capable of producing nearly verbatim copies of (parts of) those texts.
to underline this point even more, we know that MP3s (and more modern and much more efficient codecs like OPUS, after that) have been psycho-acoustically optimized to store exactly the least amount of data that will get "the point" of that music across to the listener, to the extent that they do not need the original recording any more. this is the stated goal of lossy compressed audio, after all. well, it also happens to be the (pretty much stated) goal of LLM companies, to store exactly the least amount of data that will get the point of that text to the reader. and it does tend to cause the readers to not really care about the original book any more.
having said all that, I don't mean to argue to lock it all up. I actually mean to argue that we should demand that Anthropic and Open AI release their weights data, and if anyone were to happen to break into them and steal that data, I would have exactly zero pity for that. because fair is fair.
> similarly, as has been shown (by numerous publishers and authors), LLMs are capable of producing nearly-verbatim versions of the texts they have been trained on, to a well enough quality that most people won't be able to notice the difference.
If that is true, you have a legal claim and can sue them. I doubt that’s true in the general case though.
The “does it hurt the original publisher” is a test for fair use BTW, just because you hurt the sales of someone doesn’t necessarily make it copyright infringement. That is only relevant if you try to defend using fair use (and it’s only part of the test that’s used to decide fair use).
It's a perfectly sensible interpretation of international copyright law.
I'm not sure if you're serious with the suggestion I could sue them. These are both US corporations, that justice system is pretty much in shambles in particular when it concerns corporations as big as these AI ones. You can dig your heels in the sand to defend that system, but you will also have to dig your head in the sand about why Sam Altman doesn't have a Disney "influenced" avatar, but one "inspired by" Studio Gibli.
And I'm not sure if you're familiar with the concept of "fair use" in the US as it "works" in practice, it's almost insulting, ask any music education youtuber.
Also even if it would work (which it very much doesn't), whether it "hurts the original publisher" is actually literally one of the criteria for considering something fair use or not. Look it up.
> Also even if it would work (which it very much doesn't), whether it "hurts the original publisher" is actually literally one of the criteria for considering something fair use or not. Look it up.
I don’t know why you’re repeating the stuff I just wrote like I didn’t. My point is that this is only relevant for the fair use defense and not copyright in general.
Here’s what I said:
> The “does it hurt the original publisher” is a test for fair use BTW, just because you hurt the sales of someone doesn’t necessarily make it copyright infringement. That is only relevant if you try to defend using fair use (and it’s only part of the test that’s used to decide fair use).
Well I think that is the greater question.
Is AI just an algorithm. Is human creativity just an algorithm?
Who, if anyone should own the copyright if you prompt AI to write a book?
I'm thinking more from a moral and philosophical pov, the copyright regime is broken anyway
This isn't that out there in our current scenario. These models compress our collective thought and effort. Why not make these publicly owned, all profits distributed back to us?
It's not meant to do anything about LLMs. It addresses the procurement of training data. I'm glad the courts demonstrate some basic lucidity that sadly seems to have escaped tech discussion sites some time ago.
Ideas are not protected by copyright, nor are facts. You need to have a very specific and 'creative' / 'substantial' expression of an idea for copyright to apply.
The output of an LLM can be easily be such, but usually not.
> … very specific …
That phrase is doing a lot of work. In the US, any writing is automatically protected by copyright. (This comment, for example.) Whether the author can claim infringement is a can of worms: legal costs, fair use … but your “very specific” phrasing makes it sound like there’s a prescription for exactly what is protected by copyright - there is not.
> Ideas are not protected by copyright.
The expression of the idea is, however. Same with facts. The fact that I live at a specific street address is not protected. My sentence construction explaining my specific street address is protected.
> The output of an LLM …
… is not protected, not matter its shape. The US Copyright Office has declared as much.
I'll link to a previous comment of mine: https://news.ycombinator.com/item?id=48968156
> You need to have a very specific and 'creative' / 'substantial' expression of an idea for copyright to apply. The output of an LLM can be easily be such, but usually not.
This is incomplete with current US law. You need the above (the typical copyright qualifiers) AND evidence of substantial human involvement in the creation.
Minimally directing an autonomous agent does not qualify.
Just to be clear, what you're referring to is the current US standard for whether a work is copywritable, not whether training on data and "regurgitating existing ideas" is fair-use. The latter is what the GP comment was about:
> There needs to be a royalty payment based on if the AI regurgitates existing ideas. That is probably the correct way to legislate this. If anything a human does can instantly be copied by an LLM, and then sent to all its subscribers, things need to change
Correct. I just wanted to clarify the statement that parent made, as it seems like lots of people have a misassumption about the copyrightability of autonomous in the United States.
Expect it will be clarified and/or changed by law given how much money is at stake, but the current state is what the current state is.
If I were developing key IP with agents, I'd be very careful to document my human contribution.
if they had any intention of doing things "the right way", they would have gone to every publisher individually and asked for a proper license.
This settlement has basically nothing to do with LLMs.
At least not as far as the courts are concerned. Alsup ruled [0] that feeding a book into an LLM is transformative and counts as fair use. Especially when they purchased a physical copy of the book, scanned it, and destroyed the original.
But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they planned to eventually deleted them afterwards)
Where Anthropic ran into problems is that they put all their pirated books into a big central library (file on a server), and planned to keep those copies forever. Including copies they never actually fed into the LLM (a point that seriously worked against them).
Alsup ruled this central library of pirated books was copyright infringement. And it's this "pirated central library" that Anthropic are now paying a a 1.5B settlement for, nothing else.
The fact that the pirated books were also used to train LLMs is legally irrelevant. Though... I suspect a non AI company could have negotiated a significantly smaller settlement.
[0] https://copyrightalliance.org/wp-content/uploads/2025/06/Bar...
> At least not as far as the courts are concerned. Alsup ruled [0] that feeding a book into an LLM is transformative and counts as fair use. Especially when they purchased a physical copy of the book, scanned it, and destroyed the original.
> But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they eventually deleted them afterwards)
The way I understood it, was that essentially the entire case rested on if Anthropics use was "transformative" or not. And since they literally destroyed the books (not just delete files, which would be copied), that made it transformative.
Regardless if they deleted files or not, if nothing existing was transformed, it would have been illegal. But because of the destruction of k̶n̶o̶w̶l̶e̶d̶g̶e̶ physical property, this ended up being legal.
> And since they literally destroyed the books (not just delete files, which would be copied), that made it transformative.
You have to be careful, just because the judge points a factor out as notable, doesn't mean that factor was required.
The destruction of source books makes Anthropic's fair use argument [2] especially air tight, but it would be a mistake to assume that act was required, or is what made it transformative.
In the previous google books case [1] (which this case cites), google borrowed books from libraries, scanned them, then returned them. They were not destroyed, google didn't even keep the physical copy.
Yet Google Books was ruled fair use, because it was transformative.
[1] https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....
[2] Note... This part of the ruling is still not about LLMs. This was about Anthropic's right to scan books and then keep a digital library of them.
My understanding comes from here, seems pretty clear to me but won't claim to be a lawyer of course:
> Ultimately, Judge William Alsup ruled that this destructive scanning operation qualified as fair use—but only because Anthropic had legally purchased the books first, destroyed each print copy after scanning, and kept the digital files internally rather than distributing them. The judge compared the process to “conserv[ing] space” through format conversion and found it transformative. Had Anthropic stuck to this approach from the beginning, it might have achieved the first legally sanctioned case of AI fair use. Instead, the company’s earlier piracy undermined its position.
https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...
Based on that I get the impression it's quite literally the destruction part that makes it transformative, without it, it wouldn't have been tranformative at all.
I've read through the order again. I can't find anywhere where Alsup says the destruction was required.
He cites three cases where a conversion from one format to another (without destruction of the previous version) was ruled to be fair use. Including scanning books with the google books case. (And referenced the Napster case, where a similar argument was rejected)
Then made the following comparison.
"Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company. This use was even more clearly transformative than those in Texaco, Google, and Sony Betamax (where the number of copies went up by at least one), and, of course, more transformative than those uses rejected in Napster (where the number went up by “millions” of copies shared for free with others)."
So it wasn't transformative because of the destruction. The destruction only made it "even more clearly transformative" than those other cases.
Like, how can destruction be required if there were previous cases where it wasn't?
The key legal point is not that Anthropic destroyed the books, but the key fact was that Anthropic didn't distribute the scanned copies. Alsup keeps returning to this point:
"But what matters most is whether the format change exploits anything the Copyright Act reserves to the copyright owner. Anthropic already had purchased permanent library copies (print ones). It did not create new copies to share or sell outside"
"But again, the replacement copy here was kept in the central library, not distributed"
The conclusion of that section doesn't even mention the destruction at all.
arstechnica isn't exactly wrong, the quote also mentioned "and kept the digital files internally rather than distributing them". It just put way too much emphasis on the destruction, and not enough on the lack of distribution.
The other thing that arstechnica are missing:
Antropic didn't destroy the books because they thought it would strengthen their legal argument. They destroyed the because it's a lot cheaper and faster to scan books by ripping off their bindings and feeding the stacks of loose pages into a document scanner.
> So it wasn't transformative because of the destruction
I mean, the parts of "in order to save storage space" and "The print original was destroyed. One replaced the other." again makes it clear (to me at least) that the destruction is pretty much what sticks out here that makes it "more transformative" (whatever that means) than the previous cited cases.
But yeah, agree that also "didn't distribute the scanned copies" seems to have mattered a great deal, as well as the destruction part.
Yet countless families, including old folks were ruined during untold numbers of RIAA suits because "converting to save space" is not a permissable use.
They used to go around destroying lives by the thousands after Napster was creating because of the invalidity of that argument.
It is a crime to make a CD of your MP3s and vice versa, and you cannot convert your VHS to DVD.
A billionaire does it at scale, well then saving space via format conversion is a grand, while the peons still can see their lives destroyed but with it hidden via the CCB secret panel. Two tier American Justice on full display. Bankrupty and seizure or worse for thee and billions for he. Format conversion legalized only for oligarchs, and of course, no appeal so it will only be a binding precedent on that one rich guy and nobody else. Tribe on both sides, keeping special rights for themselves that are illegal for everybody else.
I'm not aware of any cases where the RIAA sued people who ripped their own CD/DVDs/VHS for personal use.
Their MO was suing owners of internet connections which were seen sharing content on file sharing networks.
Why is scanning a book transformative(a la Google) but reading data off a CD and putting into a digital format not?
How is streaming bits of the music from your computer not transformative?
The RIAA (and the wider copyright industry) were careful to never bring a case that might rule on the issue of "ripping data from CDs and converting to digital".
What they did was bring a case against Naspter, which ruled that ripping data off CDs AND THEN sharing it to millions of people over the internet was infringement. Not because of the ripping, but because of the sharing. The RIAA then somehow managed to twist public discourse to interpet the ruling as "ripping CDs is illegal".
They were careful, because the Sony Betamax case had already ruled that recording TV of the airwaves was legal, which is already a weaker case than ripping CDs you own. They knew such a case would likely rule against them, and they found the ambiguity to be much more useful.
And later cases like the google books case, and this Anthropic one provide even more evidence that the courts would likely rule that ripping CDs was legal if such a case was ever bought. (Though, it really depends on what you do with the digital copy)
This is not backed up by any evidence. Ripping CDs was never illegal. The DMCA made the circumvention of an effective copyright protection mechanism illegal, which made ripping DVDs and Blu-rays a crime. But that's separate from copyright itself. The RIAA sued Napster users not because they were converting files, but because they were obtaining them from others without a license.
There is a lot of evidence. I lived through it. Every family with children and an internet connection or MP3 player was terrified of getting ruined suddenly via a letter. It was in the news every day about some other grandpa or single mother losing their house.
Ripping CDs was long illegal. Perhaps the Librarian of Congress made an exception. Now they hid everything behind a CCB that is like Arbitration so we will never know because they have hid almost aspects of societal justice about copyright and business labor behind arbitration style secrecy. The most useful courts are secret and now people believe there are no proceedings and they do not understand how much of our society was litigated and debated before.
Here is an article from 2008 Specifically explaining that ripping a CD to format convert for personal use is illegal and the RIAA and Sony BMG saying it merited suit but they had bigger fish to fry.
1. https://www.npr.org/transcripts/17814972
Mr. FISHER: That's right. So then, you have to ask yourself, why is the industry continuing to cling to that notion that there is no such legal right? (1)
"Bigger fish to fry" is not the same as "legal".
People selling software to easily convert VHS to hard drive were also punished. For decades they were very clear that format shifting was outlawed. But now that it supports centralizing power and creating a permanent class of info-priests to rule the society, they allow it for them.
Frankly, making all the justice system secret is why the media had to turn to personality cult nonsense for most reporting. All the great stories of the past were informed via the justice system activities. Since all the court stuff is secret now, all they had to talk about was Donald Trump.
"Discovery" provided the bulk of news facts before they secreted away all the justice system proceedings for liability, labor, negligence, medical care, copyright.
It used to be possible to know stuff about America and there was "evidence" all over the place. Now there is never any evidence for anything anywhere. That's Scalia's legacy thanks to Concepcion, absolutely gutting the ability of the society to use Hawthorne effects to discern legality and behavior.
This country used to have evidence for everything, and now a lack of evidence is so common that it is a trope level popular refrain.
You're confused about what was actually illegal and what the industry wanted you to believe was illegal. They didn't want to take anybody to court for actually ripping a CD because they didn't want to lose and have the precedent set like it was in Sony versus Betamax. This isn't "bigger fish to fry". This is "terrified of the precedent".
> The dispute arises from a suit the RIAA filed against a man in Arizona who bought CDs, copied them into his computer as MP3 files, and then put them into a shared folder that other people could access through Kazaa, a computer program for sharing music. He's being sued for that last part.
Then they state what they wish were true:
> But according to Marc Fisher, legal documents and some statements by industry officials make it clear that the industry regards the simple act of copying a CD onto your computer or your iPod as illegal.
But just because they wished it to be did not make it so. Trillion-dollar companies have provided end-users with software to rip CDs (including iTunes), and there's never been a court case over it.
it was a crime to run your own unregistered taxi service in many places until uber came and the laws changed to adapt
They did not change the laws. The rich tribe guy ignored them and they let him off just like the Anthropic guy. They did not change the laws. The old businesses just folded and the new ones via unlicensed independent contractors made cottage industries out of small scale fraud and tax evasion.
How many Uber drivers can show their local business license for every town they pick people up in? How many have sales tax accounts for their state? Every uber driver without them should have been charged with the same crime as Al Capone.
Now they are trying to control the knowledge, and the vehicle driving, etc. via AI.
It is frankly a tribe takeover via mass criminal activity.
He should have been charged with tax evasion for every pickup in a place where he lacked a business license, but the tribe would never allow it. Compliance is only for the other guys.
Only the dumb local guy graduating high school trying to earn a living has to worry about legal compliance since the rich guys are too hard to prosecute.
They did not change laws. They refuse to enforce them and we are being taken over by the reincarnated legion of Al Capone as a result.
> They did not change the laws.
yes they did?
in Québec: https://www.ctvnews.ca/montreal/article/uber-is-officially-a...
in France: loi Thévenoud and Grandguillaume (which were the follow up to negociations between the french gov' and uber), etc.
other countries are the same around the same period, e.g. https://legislation.nsw.gov.au/view/html/inforce/current/act... etc etc
Nope, that's a little bit of sloppy writing on the part of Ars. I am not a lawyer, but I'll be happy to discuss the technicalities with anybody here. I'm fairly passionate about the technicalities of copyright.
What does that mean to be "transformative", as a defense?
I thought that was explicitly disallowed use... like turning someone else's book into an audiobook and selling streaming access to it.
Or writing a film adaptation and selling the film.
Clearly I was thinking about it all wrong. Those wouldn't be allowed, even if you legally aquire the book from a store or library.
You're mistaken.
The transformativeness of the use is independent of the destruction of the books. The destruction of the books allowed them to argue that they had not duplicated them, and was instrumental in the argument supporting the legality of scanning them. But that's entirely upstream of the way the data was leveraged, which is what is critical in the argument about the use being transformative.
It's easier to ask for forgiveness than permission, right?
It seems to be the modus operandi of corporations in general: they commit any kind of infringement they want and then later they go for a settlement with a value that's, of course, not too big for a company too big to fail.
In the meantime, the average person or company gets shafted.
In my opinion, we are one step away from AI companies capturing the entirety of copyright legislation.
You do indeed appear to have a valid point. Many "chosen" companies, like Uber for example, appear to have broken numerous laws. Legal action against many such companies comes suspiciously slowly, where they have already obtained massive profits and value, before the possibility of being shut down comes. Then, when they are finally pulled into court, they have all kinds of money for the best lawyers and have already paid the right politicians (and others).
When the legal judgements for wrongdoing are finally handed out, they often come across as just an inconvenience or kind of tax, which is easily handled in comparison to the profits they've already made. Yet, if average Joe or persons not considered as being of "the right type" were to do such actions, they quickly get the full book thrown at them. Often, the full measure of legal punishment, where their company and life is or about nearly over.
Paying this sort of fee in the first place is itself regulatory capture because only the big companies will be able to pay it. If they can pirate to make an LLM then so should us commoners be able to too.
As far as I'm concerned, the courts are wrong, and training on ill gotten copyrighted material is not fair use. Given the clear value of highly trained LLMs, the investment they have taken on, and the amount of disruption to the existing economy they stand to make, in a just world, the people who created the training data deserve some level of compensation. I think, in the US, they are very afraid of falling behind China, who doesn't give a shit about intellectual property, but that doesn't mean we aren't crossing an ethical boundary, acting like them.
copyright is bullshit
Copyright is what stops someone from copy+pasting a book that took years to write, then selling it $1 cheaper than the original author on Amazon or whatever and making a margin 1 million percent higher than the original author.
Imagine a society without copyright… only physically intensive jobs could make money because everything else would be pirated, ripped-off or free. Thus, only those who are financially independent could afford to publish. Because the world really needs more rich class propaganda…
Right, it's about incentivising intellectual work. While I have big issues with the copyright system, like all the extensions lobbied for by Disney and friends, it did enable a lot of good work to happen.
> it did enable a lot of good work to happen.
How do we know that when we don't have a copy of the world without this regime? How much more and greater works could have been produced without such a repressive system?
A really successful work becomes part of the culture, and remixing, derivatives and other modes of integrating cultural artifacts are prohibited. Why should we allow corporations to own our culture?
You’re arguing that freely remixing original work will give rise to greatness that’s even better than original work?
Isn't most creative work synthesis rather than unique whole-cloth creation? Look at what happens with software when it is open sourced and allowed to be remixed freely. Are we better or worse off because of it?
There's nothing that prevents people from remixing things that are not copyrighted and create something amazing that others are interested in or of cultural value.
With open source, I should note, its remixing is in fact governed by copyright.
I think the idea is if "AI" can solve math proofs that humans haven't for a century then if "AI" freestyles stolen art and literature then it might create something as good if not better because of resources and processing power
There might be something to that logic but art and literature doesn't obey rules like math and copyright exists to protect creators
Yes? I don't understand how this is even a question, this is exactly how it worked throughout human history.
In what way is copyright preventing it today?
I can't make and sell a remix of an existing IP and even in cases of free distribution it can get dicey legally.
Possibly. Sort of like how Disney remixed basically everything from existing fairy tales for decades then made it so nobody else could.
Yeah, I can't AB test against a world without copyright at all, but I think there's sufficient evidence to believe that a lot of stuff would never have gotten done without copyright to ensure it could be done gainfully.
The importance of striking a balance between incentivising creation and enriching culture was why the original copyright term was dramatically shorter. The modern term of owners life + 80 years or whatever it is, is clearly ridiculous. 20 years before entering public domain seems pretty reasonable.
There's unfortunately also some pressure against people using legitimate public domain works. E.g. youtubers getting copyright strikes for playing public domain music because it's too similar to a specific copyrighted recording.
Go read a few fanfics and tell me you still think that theres added value.
Quite rich coming from a creator apparently, only certain creations are deemed worthy by you, seems like it invalidates your entire stance on the sanctity of art that all these pro-copyright people seem to hold.
That's oversimplifying things.
Copyright doesn't actually stop me from pirating a book or an mp3 right now. Heck, I'll just download a book right now. Bam. Done. Some things are so difficult to keep from being pirated, such a photographs, that saying the copyright system protects photographers strikes me as a bit silly. It does protect some commercial photographers if a magazine wants to sell their photo sometimes, but that's a very very small slice of all the photos in copyright that are being shared online right now.
Also there are other systems that might protect an author's financials. Off the top of my head I imagine you could do a netflix model where every citizen pays some taxes to consume intellectual property like a utility. Then the goverment finds a way to measure what is being consumed and gives each author a share based on the rate of consumption. In fact the "intellegence is a ultility" ramblings of Sam Altmen sort-of point in this direction. But that's just one idea thought up early in the morning when its too hot to sleep properly. I'm sure there are many others.
> It does protect some commercial photographers if a magazine wants to sell their photo sometimes, but that's a very very small slice of all the photos in copyright that are being shared online right now.
That is a very small slice thanks to copyrights. Without copyrights then corporations stealing from the small guy like this would be the majority of it.
> Thus, only those who are financially independent could afford to publish. Because the world really needs more rich class propaganda…
Yes, all the rich class propaganda being pushed by open source developers working on software in their free time.
How many times were hugely popular books rejected before a publisher decided they were worthy?
Copyright far more protects the wealthy than the good. They don't need to sell your book, they just need to own the book that people are buying right now. Giving your book a chance to sell would dectract from those sales.
If there were no copyright anyone trying to sell the book $1 cheaper would be undercut by someone selling $1 cheaper them them, and so on. The financial incentive to do that goes away. People then choose to distribute based on different incentives, like the fact that they have seen something worthy that others should see. We have almost completely lost that today because the financial incentive doesn't care what it is as long as you buy it. That might lead to a world dominated by an optimisation for whatever it takes to get you engaged, or worse, addicted. That world might really suck.
There needs to be a way to support the creation of art. Copyright lets a few corporations decide the subset of available art is seen enough and available to pay for (in the hope that maybe some of the patment gets to the creator). It is not a system that works in the modern world.
Well, there is nothing to distribute if the author is not incentivized to write... which you seemed to skip past.
Indeed, perhaps I should have said
There needs to be a way to support the creation of art.
Great, that way is called copyright. The author has the right to control who has the rights to distribute their work, and can require compensation in exchange for that right; what economists refer to as "selling".
People have been paid artists before copyright even existed, what some might call patronage.
And?
Therefore copyright is not necessary.
That's a big leap
>There needs to be a way to support the creation of art.
Yeah, it's called "copyright."
If money is the only incentive, then it's not a product of artistic work.
Also current copyright laws only exists to fulfill the constitutional mandate to promote the progress of science and useful arts. There are a lot of alternative ways to fulfill that mandate that don't include a lot of the baggage we have presently in copyright law which is now slowing down progress.
>If money is the only incentive, then it's not a product of artistic work.
This is just bullshit and no one said it's the only incentive.
> There are a lot of alternative ways to fulfill that mandate that don't include a lot of the baggage we have presently in copyright law which is now slowing down progress.
such as??
I can't tell if this comment is satire or not.
You're speaking to the generation of pirates. What? Suddenly everyone is hanging up their high seas hat to capture the virtue signals of current sentiment?
It's funny because copyright only benefits the rich now. Record labels hold all the copyright to songs, same with publishers for books, Disney made sure it lasts over a hundred years. The days of copyright being held by individuals in any real sense is long gone.
This reminds me of this argument with libertarians/ancaps:
A: rich people pay less % in taxes than wage workers, we should close the loopholes
B: but taxes are immoral to begin with
A: ok, but can we do something now about the unequal enforcement? Unrealized gains, tax havens, trusts, fake charities, etc?
B: well a society based on property rights… ackhully you should read this book by Mises/Rothbard/Rand
Never really understood how libertarians expect to have someone making guns for their fiefdoms when there is no one to enforce property rights for said gun elements and manufactories.
Libertarianism is not a philosophy. It's selfishness taken to extremes and trying to find ways to justify it at a societal level. The only reason we're the top species is because we're ultra social and have culture, which is inherently a social trait (don't eat those red berries, they're poisonous). Libertarianism want all the benefits of working together with no actual thought into how that working together happens in real life, including punishment for bad behavior.
Yeah its wrong in very very similar ways to communism.
What very very similar ways would that be?
I guess they presume it requires on the good will of everyone to live peacefully without violating (non-existent) property laws over, say, robbing you in your sleep.
Surprise: the comment above was downvoted in the bastion of libertarianism :-)
https://youtu.be/lh2__MN-FTU?si=LXIaljh__s8fD75l&t=1568
About 3 minutes of video worth watching.
I would be fine with abandoning copyright ... If it is done for everyone equally, and not just tech giants and VC money businesses get a free pass, while everyone else still has to follow the copyright laws. Lets go ahead and usher in an age of free information and experiencing all forms of human expression for everyone. But lets also come up with a way, to compensate our creative minds and our educators and artists. How about that UBI? We stand much to gain as humanity.
This. Copyright is a flawed system. There can be alternatives that allow more than 1 player to play and not create monopolies.
For example. I invent a new method of power washing. I start a power washing business using new tech. I file the tech for patent and copyright-equivalent use. This is then made available to other power wash companies that wish to use the tech and be certified in it so long as a small portion of their revenue goes back to the inventor for a set amount per volume, or something similar of a metric that has a cutoff after a point.
This will breed new industries, create new jobs, introduce new innovations, and allow the markets to move on from being strangled by one giant corporation.
Isn't that just...patent licensing? But I agree that it should be a forced outcome so everyone can use it rather than waiting a ridiculous 20 years.
In that case AI should just be open source/weight. I don't agree with copyright in general but I see where you're coming from.
I think this posture is hugely beneficial to China if they can commoditise the hardware.
So will you owe life long compensation for all the knowledge you got from books too? How about all the pirated books, music, movies, etc you consumed? When will you set up a life long payment plan to corporations that own these rights, because I have a bridge to sell you if you think any of this settlement will go to any of the people who created anything.
I’m guessing you have some kind of imagined idea of some small author being compensated handsomely for his book and all future earnings that could have come from it. Reality though is that between the attorneys that will run away with some high triple digit millions and the corporations that own the rights to the subject works, there will be measly “checks” for any actual person that created anything, i.e., an artist or author.
In an odd way, this whole case is really just “capitalism” cannibalizing itself, i.e., publishers greedily and also in a terrified manner trying to steal away as much capital from the technological shift to AI as possible in order to either create a buffer or fund their transformation to adapt to what AI means to the very nature of writing itself, let alone publishing.
I suspect human writing could survive, but I don’t see any room for publishers.
> So will you owe life long compensation for all the knowledge you got from books too?
No because we are people and the laws differ for people, corporations, and machines.
I think people forget that laws are perfectly capable of carving out exceptions, leaving purposeful ambiguity, expressing intent, etc. Yes, humans can have special rules, and very obviously should since laws exist to improve human lives.
I am actually not settled on either side of the matter and I have not forgotten that, but I think what we are really looking at is a rather more complex matter than people want to make it out to be. We are holding several but at the very least contradictory positions and they are incompatible.
Either it is theft or not theft. Why would you stealing from me deserve an exception, but when a group of people in an organization you may refer to as a corporation steal from me, you want them to pay me? Do as I say, not as I do or something like that?
Of course exceptions can be carved out, but they cannot be just, inherently. The problem is that we have allowed our ruling maniacs to create a fiction that organizations are people, which not only have more rights, and less responsibilities, and even less consequences/penalties; but also confers upon the individuals that make up the corporate person rather extreme super powers like being able to commit crimes up to outright murder, and there not only are effectively zero consequences for or to them but in most cases today they immensely profit from it and then shield that money from the victims seeking justice.
The underlying issue, why I am not settled on this matter, is that it is inherently contradictory because the facts and underlying assumptions are all so distorted and perverted that there is no good answer to be had and it's really just a matter of rule of power, feigning rule of law.
> So will you owe life long compensation for all the knowledge you got from books too?
You're just falling into the trap of anthropomorphizing the phrase "training" in the context of LLMs, which is not the same things as what humans do. There is no evidence they are the same thing and there is nothing to support the notion that what an LLM does when it "trains" on a book is equivalent to a human reading it.
Whatever "training" is, if you can't persuade the machine to spit substantially the same text back out verbatim, it's clearly not something that falls under copyright law either, because there's no copy.
Yes, for some texts that's possible. But for the vast majority, it is not.
Can you cite any information on this not being possible for the vast majority?
Or is it simply that the correct prompt hasn't been written for all possible cases?
I also fail to see the difference if logic/harnessing is added around a vector database that can output the complete corpus, but simply is instructed not to.
It very clearly is still compressing the information into the vector weights, and then recovering that information, thus the information is encoded.
Why is a vector database somehow completely different from maintaining a library of the text itself?
Information entropy. The amount of data an LLM ingests cannot be compressed to the size of the weights even at maximum theoretical compression.
Is that relevant? I can use a lossy compression algorithm such that the original could never be recovered from the image I've produced, but that derived image would surely be under copyright.
LLMs are obviously capable of producing "exact" phrases as well. Ask it to give you famous quotes, it can do it. Ask it to read a paper for you and cite it, it can do it.
> but that derived image would surely be under copyright.
I wouldn't bet on that. https://en.wikipedia.org/wiki/Campbell%27s_Soup_Cans
I don't know that this really challenges anything relating to compression.
Why would a derived image be under copyright?
Because that's legally the case? I don't understand the question. Using a lossy compression algorithm on an image does not remove its copyright protection.
You are asking to prove a negative. But even assuming that the model is capable of returning every bit of its training data verbatim (a mathematical impossibility) that would not be enough as mere capability is insufficient here. If capability alone were the standard any library that also has a photocopier / scanner would be in violation.
To prove distribution of copyrighted materials it would have to be practical and actually used in the wild by people to circumvent copyright and generate copies of those works. Again, I can't prove a negative, but that isn't the standard, and nobody has shown a practical exploit here.
There was a paper a while back where (from memory) they managed to coax 75% of the original text of some internet-popular books out of an LLM. Harry Potter, 1984, etc. That's why I said it was possible for some texts.
My assumption is that multiple copies in the training data "wear a deeper groove". I believe those are infringing, and should be dealt with on a case-by-case basis. But the vast majority of text doesn't wear that groove.
(Edit: Think it was this one https://arxiv.org/abs/2601.02671)
Nah, I can't prove a negative. But Common Crawl is 12 petabytes and is not the largest part of what these models get trained on. DeepSeek v4 Pro is, what, 865GB?
That's one hell of a compression ratio, if it can do what you claim.
> spitting out verbatim text
The New York Times lawsuit is resting on the point that large chunks of undigested articles can be vomited out. OpenAI tried to have the lawsuit thrown out but the courts permitted it to continue.
The Times... alleged that OpenAI's ChatGPT and Microsoft's Copilot had produced near-verbatim replicas of copyrighted articles, that the chatbots generated hallucinated content falsely attributed to the Times, ...
https://en.wikipedia.org/wiki/The_New_York_Times_v._Microsof...
> spitting out verbatim text
> had produced near-verbatim replicas
Don't get too hung up on the preciseness of the copy - the courts won't. I doubt that spitting out an existing article with a few adjectives changed would be considered transformative.
It's possible. Would be interesting to see their evidence, and to know whether they can reproduce it for arbitrary articles, not just ones that have been endlessly republished on the net.
> if you can't persuade the machine to spit substantially the same text back out verbatim
That's exactly what they've done in a number of the lawsuits, so I'm not sure why you think that hasn't occurred.
This reliably inevitable rationalization comes up in every thread it seems, and its ultimate goal is to humanize AI. This is what the big guys want us peons to believe and it works so well, I have even been lectured by an AI for being rude, the implication was that I was logged in and it would be a shame if anything happened to my account.
Quit trying to make AIs human, people who are trying to make AI human keep forgetting that humanized AI's have only the morals relevant to their mission, there is no profit in humanizing AI's because if we continue on this track of humanizing AI's, we being stupid humans will grant them civil rights expecting these new AI's with rights will somehow respect our rights and thats a fundamental misunderstanding of how AI'S actually work.
> As far as I'm concerned, the courts are wrong, and training on ill gotten copyrighted material is not fair use.
It’s important to remember that a court’s job is to apply law to a situation. When a court gets something wrong it’s a misinterpretation of the law and will, by definition, be overturnable on appeal. I suspect that your objection isn’t that the court is wrong, it’s that the law is wrong.
It's not a settled area of law and there is a SDNY judge that has a completely different application of the fair use analysis in the same exact context and came to a completely different conclusion (that it is not fair use).
What do you know, you can get someone to support any message or arguments you want.
Lesson in there about experts and politics.
I would like to see a citation on that b/c I am unaware of it. The only case I see in SDNY is the NYT v OpenAI case which has not been ruled on yet. https://www.reuters.com/legal/legalindustry/copyright-law-20...
Sorry, I'm thinking of Kadrey, where the court rejected Anthropic's "training" argument and provided an explanation as to how author litigants should demonstrate market harm in order to succeed on a fair use analysis, a factor that Alsup did not effectively weigh.
I suspect the market harm angle is not going to work out either based on the one study I know of on the topic: https://www.nber.org/papers/w34777
> We document a tripling in the number of new books coming to market between late 2022 and late 2025 that mirrors the use of AI that we detect in new books. The effects of this influx on consumer welfare depend on the quality of the additional books. The average quality of new books has fallen with the LLM-induced influx, and books with detected AI are substantially worse than human-authored books, so that much of the new work is of little value to consumers. Still, the LLM influx has delivered some books in the middle range of the usage/quality distribution, and the LLM-era entry process delivered seven percent more consumer surplus from books than the pre-LLM process in 2025.
...
Moreover, the arrival of LLMs does not appear to have displaced activity by incumbent authors. Despite the controversy surrounding LLMs, their effect on book consumers – like other cost-reducing technological changes in the cultural industries – is positive. However, because the new books are mostly of low quality, the effects are modest</i>
So not only are existing authors unharmed (because most of the new competition is slop) there is even a small improvement for consumers.
Yes, ultimately the problem is that the law is vague or inadequate. The courts have their definitions of fair use, which are their best efforts at interpreting the law, and I have mine, which is different.
William Roper: "So, now you give the Devil the benefit of law!"
Sir Thomas More: "Yes! What would you do? Cut a great road through the law to get after the Devil?"
William Roper: "Yes, I’d cut down every law in England to do that!"
Sir Thomas More: "Oh? And when the last law was down, and the Devil turned ’round on you, where would you hide, Roper, the laws all being flat? This country is planted thick with laws, from coast to coast, Man’s laws, not God’s! And if you cut them down, and you’re just the man to do it, do you really think you could stand upright in the winds that would blow then? Yes, I’d give the Devil benefit of law, for my own safety’s sake!"
This is why the idea of being "Vogelfrei" or "lawless" was honestly a terrifying concept in the middle ages. They are neither bound by law, nor protected by law.
A lawless man can be struck down with force without persecution by law, because they are lawless.
"I haven't loaded an advertisement in 20 years, I have 6TB of movies, 2TB of music, and seemingly endless file trees of mangas, all acquired for free over the years. Now having not said that, I beg you enforce copyright on these AI labs, so I can get a cut of their revenue for my years of writing well researched comments on the internet"
The internet, in true internet fashion, still has the general logic level of a 15 year old.
IMO, using copyrighted works to train models should only be "fair use", if the models are then released as (at least) open weight, so that the public can benefit from it. (Although as noted by a sibling, this would require a law change, not action by the court).
I’m not sure if that’s enough but it would be a great start.
Adobe’s ereaders had a disclaimer that their books cannot be read aloud. There’s clearly precedent that this sort of transformation was disallowed by publishers at the time. Interestingly, at least the audiobook of the latest dungeon crawler Carl has a disclaimer that it can’t be used to train AI
That's a good ruling, because otherwise only the big companies can afford to pay for enough content to make an LLM (say goodbye to open weight or research LLMs). Having a fee like this is actually a form of regulatory capture.
> But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they planned to eventually deleted them afterwards)
The court says otherwise.
> Such piracy of otherwise available copies is inherently, irredeemably infringing even if the pirated copies are immediately used for the transformative use and immediately discarded.
Then it says it doesn't need to decide on that basis because they kept it not just for training LLMs, but also for building a central library. Which seems a bit ridiculous, because the sole purpose of the central library is to train LLMs.
Who would have thought, that this is the way, which we take to arrive at the burning books stage again? They neatly line up with historical perpetrators in that regard.
>But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they planned to eventually deleted them afterwards)
You're not. Even if training is fair use, it doesn't mean you can steal copies to train the model. It just means the training itself isn't an infringement (in Alsup's opinion). Stealing the copies of the books was an infringement and that's exactly the liability that Anthropic settled.
Hey I just came up with this idea, I'm going to feed copyrighted books into my LLM that remembers them verbatim, and then people pay me to ask the LLM for complete copies of a book.
Wait, no, not verbatim. It transforms upper case into lower case and vice versa.
If it is based on all our data, we should all own it and democratically chose what is done with it or profits it generates
The correct way to legislate this is to abolish copyright. It is strictly a negative force. Nobody makes art because of copyright, only in spite of it.
Commercial enterprises make content, like Marvel movies and Netflix series, because of copyright.
Commercial enterprises stand on the shoulders of lax copyright laws. For example, foundational Disney works would have been illegal for them to make under the copyright laws they have since purchased.
https://drewdevault.com/blog/Alice-in-Wonderland/
Copyright is a textbook ladder pull.
What copyright law helps them? They have the least to worry about copyright as even if someone copies the movie script or something it's not like their views will be gone because of that.
I am not against trademark. e.g. Disney has a right on who can sell Mickey Mouse figurine, or Marvel has right over Iron man character and franchise.
People make art to also get recognized for that art. Otherwise they would keep that art secret at home.
Without copyright, anyone can copy the art and call it their own. What is then the incentive for the creator to share the art, if there is neither monetory gain and nor fame. And worse than them being recognized, they might even get accused of copying their own art if someone else became famous due to a copy.
Society would miss out a lot.
There are lots of famous artists that are older than copyright.
> And worse than them being recognized, they might even get accused of copying their own art if someone else became famous due to a copy.
You mean like right now? Here is A24 claiming copyright for Backrooms related media that came out before their Backrooms related film: https://kotaku.com/backrooms-a24-copyright-strikes-kane-pars...
Being "accused" of copying literally does not matter if copyright didn't exist. This framing is only an issue under copyright.
> Society would miss out a lot.
Society actively misses out a lot. We could have had tons of derivative art that has been buried for the sake of propping up companies. We could have had Aaron Swartz. Abolish copyright.
> There are lots of famous artists that are older than copyright.
There are lots of famous artists that made art before color image capture and reproduction (1930s-60s) and digital image capture and reproduction (90s-00s).
Copyrightless artist fame and economic viability is enabled by a lack of widely-accessible, cheap reproductive methods.
> You mean like right now? Here is A24 claiming copyright for Backrooms related media...
Since resolved: https://kotaku.com/backrooms-director-kane-parsons-a24-copyr...
Turns out when you outsource copyright policing to minimum wage folks, ambiguity goes out the window.
> Society actively misses out a lot. We could have had tons of derivative art that has been buried for the sake of propping up companies. We could have had Aaron Swartz. Abolish copyright.
The problem with absolutist arguments is that they ignore inconvenient facts.
At a time when creative and art economics is under siege, how would copyrightless art make enough money for the creators?
The fact of the acquisition of large swaths of copyright rights by large corporations does not negate the fact that artists need food (and ideally, a place to live and a way to provide for their family).
There are positions that might enable that (Hey, what if we banned the assignment of copyright to corporations? Human only? Original creator only?), but none of them are stripping all rights from intellectual property.
Many of the most valuable paintings in the world are out of copyright. It has not diminished their value, because people still value originality even if it's not enforced by the law.
> And worse than them being recognized, they might even get accused of copying their own art if someone else became famous due to a copy.
This happens now all the time, and the winner is determined by who can afford the best lawyers.
What if copyrights are shorter, but vigorously defended (other than fair use provisions)?
Can we use the modern tech (AI) to policy copyright infringement, to liberate the culture and business?
The shorter the copyright the better. Zero is the best.
> There needs to be a royalty payment based on if the AI regurgitates existing ideas
This does not do enough to fix the root problem.
People who live right now, who happen to have written or produced anything that AI works with, build on the back of humanities combined knowledge, will become outsized beneficiaries of AI, with the AI wave offering new ways of monetizing their work – while everyone who has not, won't be.
It's simply not good enough. We have to make sure people broadly benefit first and foremost.
It's easy to beat up on OpenAI and Anthropic, because they have lots of money and knowingly broke the law, but writing the book was onetime work too. Do we really want to turn everything into recurring revenue stream to skim of? How would that even work for an open weights model? Would you say the same about a human educating themselves from a book? The answer has to be more than pearl clutching for poor starving little authors (and the not so poor class action lawyers).
I agree. If I pirate a book and share it on the web, and I get busted for doing so, and subsequently pay a fine, I don't get to KEEP sharing it on the web.
Now, if I license the book, I might be able to come to an agreement with the author/publisher whereby I can share some of it.
The post specifically proposes royalties for ideas from books, not royalties for the books themselves. You would absolutely still be able to share ideas you learned from the books you pirated in that situation. It'd be insanely draconian if you couldn't.
(Then again, US copyright law often is insanely draconian.)
> royalty payment based on if the AI regurgitates existing ideas
That doesn’t make sense. You cannot copyright an idea, only the specific expression of the idea.
There needs to be a royalty payment based on if the AI regurgitates existing ideas.
We're trying to own ideas now?
> There needs to be a royalty payment based on if the AI regurgitates existing ideas.
The settlement does not pertain to any outputs
I'd rather see AI studios do the same as the film industry, pay a one-time up front cost per major model (or major.minor?) depending on how they contract it out. This also allows smaller startups to license books for less than a larger frontier studio would. In theory and hopefully, the pricing would not be too insane per book, you want them to rent more books and spend more, not go back to pirating right?
how would you do that? you can copyright words, but you can't copyright an idea (you can patent some ideas, but not all of them)
If you have the time, read the judge's response to the motion:
https://storage.courtlistener.com/recap/gov.uscourts.cand.43...
The big deal for publishers and authors is the payout per eligible title is $3k. For a traditional publishing contract involving one author, the amount will be split down the middle.
The other thing which caught my eye is the judge slashed the class counsel's fee by half, from 12.5% ($187.5m) to 6.8% ($101m). The class counsel's unreimbursed litigation expenses were $2.6m.
The three class representatives get just $15k each.
Realtors are capped in the percentage they can take for selling properties. Brokers and financial advisers are capped in their fees. My presumption is that the only reason this very standard and reasonable regulatory pattern doesn’t affect lawyers is because they tend to be the ones writing and enforcing the regulations in the first place.
There is no legal maximum for the percentage real estate agents can take, in the US. Rates are also not fixed, by law, and are required to be negotiable. There's a general standard for rates (typically 5-6%, split between agents/brokers), but there's nothing stopping them from setting it to 99%, other than the fact that people won't pay it.
Source: was a licensed real estate agent for a long time.
I might be misremembering, as I worked on the loan side, but wasn't the 6% standard set by the state's realtor association until 2024?
No. "Realtor" is a trademarked term for a member of the National Association of Realtors. Real estate agents are licensed by state governments, but prices for real estate agents are not legislated by state governments.
If this is the 2024 settlement that you are referring to, it did not say anything about the price a Realtor can charge:
https://en.wikipedia.org/wiki/Burnett_v._National_Associatio...
>The cooperative compensation rule has been eliminated as a result of the settlement. Seller's agents are no longer required to offer compensation to buyer's agents when listing a home for sale on a Realtor-owned multiple listing service. In addition, Realtors acting as buyer's agents must enter into contracts with buyers before touring any homes, allowing buyers to negotiate how much they will pay their buyer's agent.
I'm aware that prices for agents are not legislated by state governments, but prior to 2024, as a realtor that was a member of your state's association, you were standardized at a 6% rate. Because that's what your association standardized. That's partially what the suit was about. What percentage of real estate transactions are done with agents who aren't members of XAR?
> pattern doesn’t affect lawyers is because they tend to be the ones writing ... the regulations in the first place
In my state, the people hired to write new bills (if passed becoming statute) have to pass their JD (law degree). Legislators only get to request new bills, they can't hand a proposal (which may have been written by a lobbying agency like ALEC or Heritage) into the system.
>The class counsel's unreimbursed litigation expenses were $2.6m.
In what sane state does it even get that high?
I've seen some YouTubes where lawyers were complaining about high bill rates and showing actual bills. One large firm billed the senior lawyers at $2500/hour and even the paralegals were billed at $600/hr.
That could pay for a lot of tokens.
Associates are billing over $1000/hr and can bill 12 hours a day easily enough. If you have 5 associates billing an average of 50 hours a week then in a month thats a million dollar bill just for the junior lawyers.
Judge Alsup issued the original order that determined they were liable for piracy but that training LLMs on books was fair use. It's worth reading if you're interested in the topic. https://www.courtlistener.com/docket/69058235/231/bartz-v-an...
Alsup is an interesting judge. He has handled several important tech cases, such as Oracle v Google, and Waymo v Uber.
He's also a longtime hobbyist programmer working in BASIC, much of it in support of his ham radio hobby. Screenshots of his shortwave propagation prediction program here [1].
[1] https://www.theverge.com/2017/10/19/16503076/oracle-vs-googl...
He learned Java to understand the Oracle v. Google case better.
His middle name is Haskell.
And his name, if you need glasses like me, looks like AI slop.
And his favorite food is curry.
He was one of the few judges that understood tech. Unfortunately he retired last year.
Yes, that is interesting. It sounds like he was aware of the theoretical possibility of a book being regurgitated verbatim. Do you know if he was aware it had been done? https://news.ycombinator.com/item?id=49000742
If he was not aware, I wonder if he still would have described the process as "exceedingly transformative" had he been aware.
Note that they're testing for 100-word passages. This is a level of memorization that avid readers can credibly also reach.
Note also that Sonnet 3.7 had to be jailbroken.
Note also that they got high memorization for a few books that were widely quoted. The books in question can probably also be "retrieved" by putting phrase prefixes into Google, which is probably why Sonnet 3.7 knows them with the precision of a fanboy. Material being widely repeated in the training set is a well-known cause of memorization.
No "avid reader" could recall anywhere near that much text. That takes dedicated effort to commit to memory. Copyright was never meant to stop people copying books anyway, it was meant to stop machines (ie. printing presses) copying them.
Edit: Apologies, I misread it as "100 pages". My point about copyright still stands, though.
I disagree that avid readers cannot complete entire passages from books they've read several times when fed a prefix.
And can these avid readers publish these recited passages without infringing copyright?
I mean, the debate would then turn on whether publishing the passages and publishing the model is the same sort of thing. I think there's mainly two views: "we know the passages are in there, so publishing the model is publishing the passages is copyright violation", and "nothing happens until you go through considerable effort to elicit the passages, so the user is committing copyright violation using the model as a tool."
Personally I think our legal system is just not set up for a world where we can download mindstates in numeric form. Would a sufficiently detailed recording of my brain violate copyright? If simulated, it could certainly be elicited to commit violations.
edit: At any rate, Anthropic are not publishing the Sonnet 3.7 weights.
I used to use the initial letters of a whole paragraph from Lord of The Rings as my password.
Some of us have a good enough memory.
So, continuing to profit--forever--from someone's else work, at scale, without their prior consent, is fair use?
It's funny that crimes can be settled in cash. IOW, everything has a price; and the price is always right. Settlement ought to be the euphemism for blood money.
In addition to the settlement, what I'd consider fair is to have these companies pay royalties in perpetuity. Of course, that's not tractable.
Yeah, I feel like penalties here should be something like 10% of revenue in perpetuity. Then companies might think twice about asking forgiveness instead of permission.
Why do they need prior consent? What sort of rent seeking do you want?
Are you saying that if someone writes a book or records a song, anyone should be able to use it for anything forever without consent?
How does intellectual work get funded in this insane world if yours, pray tell?
You assume your premise. But plenty of "intellectual work" is already done without legal cover. It just typically attracts normal profits, rather than super-normal rent-seeking ones.
I, too, hate rentseeking. Owning one's intellectual output, however, is not in and of itself rentseeking.
Rents are just amounts beyond what's needed to cause the thing to exist. At the point of copying something, it already exists.
So payments for the right to do so aren't payments required to bring anything new into existence at that point, save for the legal fiction.
Now you might argue that the future copy-licencing rents are necessary to bring the _original_ creation into being. But that doesn't make them _not rents_.
But I would say that's the second assumption you're baking in here.
As in, we live in a world where e.g. the movie Toy Story exists. Now, certainly Toy Story does provide some good or value to the world. But I don't think you can assume such things provide more value than e.g. open science, free transformation of works, etc.
I get that people enjoy our current IP culture but saying certain things wouldn't exist in an IP-free world is just an argument from consequences that doesn't even really compare consequences between the two.
Intellectual property protections have a finite lifespan to begin with, and as a tenet of Western Civilization are barely 300 years old.
For works published before copyright laws existed or after property protections expire, anyone should be able to use it for anything forever without consent.
Intellectual work still manages to get funded in this 'insane world' - although given the classical artist/patron system has given way to state-based grants and a select capitalisation of Art post-Warhol, the concept of Universal Basic Income tends to be promoted the desired successor.
Speaking of insane worlds, how does the concept of the Public Domain work in yours?
Are you really saying that because IP protections are (and should be, I agree) time-limited, they're not doing anything in the first place? This is patently ridiculous.
> Speaking of insane worlds, how does the concept of the Public Domain work in yours?
Can you elaborate on what you're asking? I don't understand your question.
Your contention was that it was an insane position that anyone should be able to use a piece of literature or a song for anything forever without consent. I simply highlighted the absurdity of that based on the fact that:
1. Copyright protections as a concept are an incredibly modern phenomenon, mostly limited in practice to Western Capitalist Democracies. 2. Outside of a short monetisable window (albeit one extended and irrevocably marred by Disney/Sonny Bono) your 'insane' hypothesis is in fact the status quo 3. Much intellectual work is published into the Public Domain, and all copyrighted work eventually ends up in the Public Domain. Your position appears to presuppose a world without such an entity.
As to what copyright actually achieves? It's mostly a mechanism by which the media gatekeepers and owners of capital use legislative and social imbalance of power to deny artist the rights and royalties for mechanical reproduction and otherwise impose financial serfdom.
This is achieved mainly by Copyright Enclosure, whereby musicians are typically pressured or contractually obligated to surrender their master recordings and intellectual property, and by contractual clauses like Controlled Composition Clauses, whereby Labels reduce the mechanical royalties they pay to artists who write their own songs, often paying below the standard statutory rate.
Lots of ways. Selling author signings, talks, authorized copies, subscriptions/merch, sponsorships. How do newfangled "content creators" fund their work? We already live in this world.
Maybe not all creative works are deserving of monopoly profits just by sitting on the ass in any case, and should stand on their own merits by producing downstream value that can be sold for whatever they can be sold for, by whoever puts in the work to deliver the value to the end user in a competitive manner. You know, open markets.
Attribution I can see. Consent or payment beyond market value, why? Just because you put in a billion hours to make a shitty $1 value output I should pay you a billion hours worth of labor?
> Selling author signings, talks, authorized copies, subscriptions/merch, sponsorships.
Most of these turn intellectual work into that of indie musicians, or outright beggars. You are stepping dangerously close to stripping people rights in favour of giant AI companies.
> How do newfangled "content creators" fund their work? We already live in this world.
"Content creators" heavily rely on IP protections. You could always try taking some youtube videos with 100M views, altering them a bit an using them as your own and seeing how that goes down. Do let me know!
> Maybe not all creative works are deserving of monopoly profits just by sitting on the ass in any case, and should stand on their own merits by producing downstream value that can be sold for whatever they can be sold for, by whoever puts in the work to deliver the value to the end user in a competitive manner. You know, open markets.
And open market is not one where I can say that you are just sitting on your ass, so I'll take your stuff and sell it.
> Attribution I can see. Consent or payment beyond market value, why?
Because it's my stuff of course! And why should one even pay market value in your world? Why not always 0?
> Just because you put in a billion hours to make a shitty $1 value output I should pay you a billion hours worth of labor?
What in the world are you on about? If the price someone puts on their IP seems too high to you, you should not pay that price. We wholeheartedly agree. Where we disagree is where you from this conclude that you can just choose the price yourself and take it anyway!
> So, continuing to profit--forever--from someone's else work, at scale, without their prior consent, is fair use?
No, that's what they got in trouble for - a lack of consent.
If the author consents, it would have been fine. If they bought the books, then it is fine. Digitisation through destruction, like most book scanning systems. As long as the original work is destroyed during the process, and you actually paid for it, then it is fair use.
If it regurgitates, then the author can sue you again. So you are incentivised to make damn sure it doesn't. That's not covered by fair use.
Its only if the original cannot be accessed anymore, and you paid to get the original. Both must be true, for fair use to hold.
That’s what the got a _slap on the wrist for_. 1.5 Billion of a payout to effectively cement themselves as one of the only orgs that can ever create one of these models because the ladder is pulled up behind them.
So, they can bu ya book and format shift it, but when I do it, it's piracy? Looking at all DMCA/DRM systems.
Format shifting has a DMCA carveout. It is 100% allowed. Since around 2000, the rule has permitted it for:
> Literary works, including computer programs and databases, protected by access control mechanisms that fail to permit access because of malfunction, damage, or obsoleteness.
DRM being covered under other laws, and being gross, still applies. And still applies to industry giants, too. Which is why most who do this, like Google, actually buy physical copies and scan it destructively, so they don't have to deal with it.
Yes, it's legal to format-shift DRM protected media, but it is not lawful. Someone has to break the law to provide me with a decryption tool.
If you asked the right politician when all these rules were being written, the intent was that each person who needs to format-shift their media would independently write their own decryption tools, use them for lawful purposes only, and then dutifully delete them the moment they were no longer needed. This is, of course, laughable.
Of course, if Anthropic was, say, buying and decrypting Kindle books TODAY; they probably could get Claude to vibe-code a DRM decryption tool[0]. That would actually be within the bounds of this asinine law. If Anthropic started off by doing this, however, they probably would have just used a decryption tool found on the Internet, and that would have invited different legal challenges. Like, is it legal to use an unlawful tool to accomplish something otherwise legally protected? The courts so far have been very hostile to ANY attempt to tie the anticircumvention provisions of the DMCA to fair use. They could easily say "No, you only get to format shift with your own tools".
[0] Related note: I really wish I had Mythos access, just so I could jailbreak my iPad on modern iPadOS. No other reason.
I dunno, ever used a thing you learned from a textbook in your job? Did you have to continue paying for the copy of that knowledge speed on your brain? No, because that's not what copyright is about.
Learning from and building on previous work is civilization. Copyright maximalism is a plague.
Yes but does the computer actually learn? Is the computer a person that read a book and remembered, a part and used that to create a novel idea or is it just cioy pasting the answer and then reselling that
A computer not, but a LLM model does learn. The original text in no way exists 'in the model', and the model does not copy-paste it and resell the original text. Not the same, but reasonably similar to a human.
We may need some new legislation. An LLM is not a person, but its also not just a storage solution.
> I dunno, ever used a thing you learned from a textbook in your job? Did you have to continue paying for the copy of that knowledge speed on your brain?
Those regulations and principles are for humans.
Either the major LLMs are software tools deployed by ostensibly-profit-seeking companies, and regulations based on the notion that "making humans pay to make use of the things they've learned is profoundly antisocial" don't apply, or the LLM companies have a bigass swarm of unpaid -er- "servants", and labor laws and other human rights regulations do apply.
> I dunno, ever used a thing you learned from a textbook in your job? Did you have to continue paying for the copy of that knowledge speed on your brain?
They are, presumably, human. We can perfectly well say that humans have certain rights without needing to give machines those same rights.
For example, we've more or less all agreed that it's fine for a human to watch a movie and enjoy the memories forever, and be inspired by it forever. But we've also more or less all agreed that that doesn't mean that a human can use a machine to record that movie and keep it forever.
> Learning from and building on previous work is civilization. Copyright maximalism is a plague.
The debate has existed for several generations at this point. You may disagree with the mainstream opinion, but it's disingenuous to frame it as "copyright maximalism".
If you buy the movie you most definitely have a copy that some machine made and you can keep it forever.
Yes. But you can't redistribute it. Which is what the AI companies are doing.
Well, as not every book is a textbook, I'd say quite a lot of what I've read never went to any kind of knowledge in my head at all. But I reckon the author still deserves to eat.
You too deserve to eat. Does it mean we're all supposed to pay you in perpetuity for that HN comment you just posted, because we read it and it's encoded in our brains now?
No... But as we're discussing people who tried to pay nothing... Maybe they should have just bought the books in the first place. And comparing my single sentence on HN, to the hundreds of millions ingested, does suggest that maybe scale changes something.
Like most things said on social media, not being copyrightable.
Does buying the book not count? Why in perpetuity? Even patents expire after 17 years.
Am I an LLM?
Why does that matter?
> So, continuing to profit--forever--from someone's else work, at scale, without their prior consent, is fair use?
Yeah, you are right. Have you been paying your dues to the authors of your math books in first 4 grades? I think 15% of your wages as an engineer would suffice. These kids continue to profit for years, and they are so many. Gotta pot a stop to that IP theft.
In my mistake I thought copyright was about copying rights, not paying for using ideas themselves. If just being downstream from a copyrighted work is infringement even without substantial similarity, then it's more like patents that expire in lifetime + a million years.
once you read the book, and learn from it and use the knowledge in your livelihood - you have to perpetually pay the author? Doesn't sound right.
Why is everyone talking about human analogies when LLMs are not humans?
Because the process is the same. When you read a book you don‘t save it as a brain file, you form memories from it. Some people can recall verbatim bits here and there but I have never met someone regurgitating a book word for word. And I‘m pretty sure I can not ask ChatGPT to output the first chapter of Moby Dick word for word. I think that would be, rightfully, considered copyright infringement.
Why is it different other than, "just cause?" No one seems to have actual reasoning to back it up while it feels very similar the other way around, is human brains and neural nets (notwithstanding that they're both called neurons) seem to learn similarly and can act on similar classes of problems like language and mathematics.
are you asking what the diff is between a human and an LLM?
LLMs are just tools wielded by humans. There's always a human behind whatever is done.
I asked why, not what.
What's the difference between humans and LLMs? The difference is that the creators of laws are humans, and the purpose of laws is for humans. Authors write books with the expectation that they will be read by humans, and copyright law was written with unstated assumptions, like the fact that books have an effect on a person's mind after reading it.
IMO the spirit of the law would prohibit LLMs from training, and the letter of the law leaves room for that only because nobody thought to write down "books are for people to read".
Laws are not only for humans, there are laws for bots as well, like anti spam, but that's besides the point because in reality behind LLMs there are humans and so humans still control them, therefore laws target them too, and now it looks like humans using LLMs to train via ingestion of books is deemed fair use.
I agree that laws don't only apply to humans, but they are only for humans. The legal system is for our benefit.
> therefore laws target them too, and now it looks like humans using LLMs to train via ingestion of books is deemed fair use
Which is orthogonal to my point.
> So, continuing to profit--forever--from someone's else work, at scale, without their prior consent, is fair use?
Yes, that is the entire history of humanity. People steal the last generations works and build something from it to make it their own.
I do recall that a "famous redditor" was driven to suicide for making works available and he wasn't even making money for it.
He wasn’t just some famous redditor. Aaron Swartz helped create Reddit and invented RSS.
If I put my conspiracy theory hat one and I always get piled on for this theory in other online communities but I think it could be possible. The theory is I think Aaron found some very dark stuff while exploring the MIT private networks, things that he was not supposed to see and could be very damaging to a lot people if they were exposed. The infamous Jeffery Epstein was donating a lot of money to MIT and its Media Labs. I think there is a much deeper story at play that the mainstream narrative is hiding with a “suicide”.
understatement is a kind of humor, when he said 'redditor'
If he found fucked up stuff regarding crimes way worse than his on university networks don't you think he could have gotten out of the sentencing entirely by cooperating against them
Epstein was relatively restrained even in his personal email, I doubt he was using MIT administrated systems to facilitate a pedo ring
It's not that simple. If he did find dark stuff, and he got caught - he will have both the state murderers and the filth on his case. That's the best case scenario assuming the state isn't on the side of the filth by which case he was cooked no matter what he did. Filth does not even have come after you. They can just send you a family portrait and you'd know the hole is deep and only one way out.
Have you seen how the legal system is protecting everyone associated with those crimes, despite how high profile they are?
Aaron Swartz founded Reddit in the same way that Elon Musk founded Tesla, but without the money.
Right with his tech skills instead of money.
No, he was added later to an already existing project and claimed credit for it.
Ok Spez
No Aaron joined reddit through a merger of the company he co-founded, where he was a developer. Musk just threw money at it
I'm not from the USA so my views are obviously biased by this. I very much doubt Epstein committed suicide, and it's been wild to watch you guys deal with it, not releasing the files, holding nobody accountable and so on. That being said, I also don't think there needs to be a big dark secret to explain why someone caught in the USA justice system would commit suicide.
It's crazy that you could face 35 years in jail for trying to free knowledge in a harmless manner. The longest anyone has been imprisoned for in my country in modern times is 26 years, 11 months and 6 days. We have a few people posed to break that record. Peter Lundin has been in prison for 25ish years, and him and Peter Madsen (the discount elon musk turned murderer who killed some poor journalist in is selfmade submarine) are contenders to people who will probably go beyond 35 years.
Not that our system is perfect. I think we're far too lenient on some crimes, but risking 35 years in jail for downloading and sharing academic knowledge... That's objectively evil.
That's the joke, why do you think they put it in quotes?
Your theory is that he found "very dark stuff" that is accessible to any MIT student? Schwartz didn't "hack" anything, he connected to their student network and downloaded articles that they had access to but the public didn't.
But why not jail like Kim Dotcom? And why no Feds jumping on Dario’s window? They are not only pirating, they also resell it!
Kim built something designed to help everyone pirate stuff. Anthropic pirated specific content.
Anthropic pirated that content to help everyone do the same.
Please post the prompts that will reproduce the pirated works verbatim. Or even halfway.
A shitty cam rip of a movie is still punishable as infringement.
That said, it's been done: https://arxiv.org/abs/2601.02671
> In some cases, jailbroken Claude 3.7 Sonnet outputs entire books near-verbatim (e.g., nv-recall=95.8%).
Subtitle translation is punishable as well…
Infringement is a civil lawsuit, with financial punishment only.
Wrong. Purple go to jail for publishing shitty cam rips. Did you read the relevant law?
I use this prompt regularly for benchmarking token rate:
I like to use hamlet. Most of them will output the first pages without issue. I tried a newer copyrighted work ("The Ones Who Walk Away From Omelas") for demonstration with Deepseek V4 flash:
Does that still work for copyrighted things in major-lab models, or only open models?
It worked for Gemini flash too, but the monitor terminated the output before I got to read more than a paragraph.
That's actually pretty cool. You're the first person that's actually showed verbatim reproduction in a discussion like this.
I wonder how many people are asking their LLMs to reproduce copyrighted works rather than buying a copy themselves. Or, more realistically, just going to Anna's Archive.
> Please post the prompts that will reproduce the pirated works verbatim.
You don't need to reproduce anything verbatim: a 1/4 resolution copy of a movie is still infringement even though it's only a quarter of the size.
1/4 resolution, but still 100% of the movie. There's not really an equivalent for a book.
I doubt many people are asking anthropic to output Harry Potter for them. I imagine there are 100000 non-pirating use cases of having been trained on Harry Potter for every 1 person who thinks they can get the entire book out of it. Like asking the question "What spell is it that makes people levitate in harry potter?" and things like that which are not infringing on any copyrights
If Anthopic had bought all the books it had trained for say at market rate we’d be having a different conversation now. Anthropic, through this settlement, has been forced to pay back, at least something… Kim would likely not have had enough money to compensate the victims and probably caused some more direct dammage by sharing pirated content. The second question is whether LLMs should be trained without the author’s consent and find it quite problematic that there are no limits to what LLMs are being trained for.
anthropic etal would not have a product to sell without their violation.. kdc had a service that just happened to be popular for pirating... how are the two even remotely similar?
the company is valued at basically 1000x the settlement it is a rounding error for them
copyright infringement was enough to get judgements that ruined entire lives when i was in my late teens and early 20s
now you get to be a founder of a trillion dollar business by extremely large copyright infringement
fuck these ghouls fuck LLMs and fuck the waste of money for this shit
> Anthropic, through this settlement, has been forced to pay back, at least something… Kim would likely not have had enough money to compensate the victims and probably caused some more direct dammage by sharing pirated content.
You're thinking civil. They're talking criminal. Criminal law enforcement does not (well, isn't supposed to) look at your ability to compensate before deciding what to charge you with.
Since we're on the criminal side - what criminal statute would apply to Anthropic? And what criminal statutes were used for the cases we're supposed to compare to?
Presumably some of the ones that applied to Aaron Swartz. [1]
[1] https://en.wikipedia.org/wiki/Aaron_Swartz#Arrest_and_prosec...
Fuck off! what about Aaron Swartz ? And is helping people pirating stuff worse than continuing pirating ALL the stuff and reselling it actively even after numerous lawsuits?
Some of you really don’t deserve good things. You should be blocked from using AI on more than one device without paying an additional subscription plan.
Aaron Swartz case was particularly egregious. If a just God does exist, that prosecutor will burn in hell.
There's no God. If we want justice we have to obtain it ourselves, in this life.
I think this is the thing that bothers me most. Seeing that 2005 YC photo with Altman and Swartz and realizing where we're at now.
Related: https://flaminghydra.com/sam-altman-and-aaron-swartz-saw-the...
As soon as we start conditioning ethics, we give up and undermine the principles behind those ethics.
- @Nevermark
Kim built something designed to share files. Are you saying Microsoft should go to jail for SMB?
You're confusing a technology for a service.
Any way, the internal emails are available where you can see the executives of megaupload knew exactly what megaupload was being used for, and even used it themselves to pirate content, and went out of their way to allow copyrighted content to remain up after takedown notices were sent.
Yeah. These AI settlements make such a mockery of past copyright enforcement victims that it's straight up offensive.
Police descended upon Kim Dotcom like he was a terrorist or something. They rappelled down helicopters and stormed his home like he was bin Laden.
Then these big techs come along and they make some absurd cost of doing business settlement.
Let's place the rage in the rage in the right place. What they did to past infringements is wrong. What they're doing now is also wrong, but less wrong.
Maybe if they could have fined him $1.5B he would have gotten away with it too.
What was Sean Parker sued for again?
Kill a man you’re a murderer. Kill a thousand, you’re a conqueror.
Kill them all... Oooooaaahhhhh you're a goddd
I harass the sea with my tiny boat and am called a pirate, you do it with a great fleet and are called a king.
This is literally the tale of Alexander and the pirate, told in “The City of God” (V century). Nothing new under the sun, unfortunately.
It was a class action law suit. I may be wrong, but I don’t think jail is ever an option in a civil suit.
right, the thread you responded to is asking why there hasn't been a parallel criminal suit
Maybe because Anthropic isn't actively facilitating copyright infringement?..
How is not? Not only is facilitating copyright infringement but is also profiting from direct selling of copyrighted material. “Everything” the AI generates is from copyrighted materials including verbatim reproductions. Sora was even more obvious.
The answer is because US government finds AI more useful than some hosting service for pirated media. It's kinda boring, but it's as simple as that.
Because no one wants that. They’re offering incredible value for what they took.
I mean.. it also sends the message that you can ignore the law if you're rich. $1.5B is like a single failed training run for Anthropic. They burn that in a long weekend because somebody forgot to abort a hyper parameter search.
Obviously exaggerating.. but not by much.
I mean, you can mostly. That's apparent all over. Money buys freedom.
It’s like you can ignore the law if you have a great idea that works out. Lots of people have ended up doing it. Uber did it for a long time. Musk, did it with the sale of Tesla cars. There are a bunch of examples from outside of the US as well.
That sounds like an oligarchy. Especially if "great" is measured in dollars rather than public good.
1. Some people want that, with good reason.
2. There's incredible value in what they stole.
3. IANAL, but I don't believe "but now everyone can write like a terrible version of the writer we fleeced" is a valid legal defense.
> stole
Copying is not theft.
AFAICT it is if you're not a trillion dollar company.
https://youtu.be/ALZZx1xmAzg?si=ugquA7uKT3ABGdws
Aaron Swartz also intended to offer incredible value for what he took, AND for no personal profit.
That’s the difference between breaking the law as an individual and doing it as a corporation.
... and I don't think people wanted him chased for it either?
Abolish copyright then. The selective enforcement needs to stop.
This is my point too. This fucking double standard/class discrimination pisses me off.
Yeah, man, sure they cleaned out the vault, but the valuable insights gained for how to keep it from happening again are priceless!
I want it to happen again. Copyright is important but I want someone that does tremendous good to be able to fall into a grey area where they’re given a free pass. But only on a case by case basis. Keep the lines fuzzy. That way we get to defend copyright but someone extraordinary also has a ray of hope of getting away with subverting it.
> Keep the lines fuzzy. That way we get to defend copyright but someone extraordinary also has a ray of hope of getting away with subverting it.
That's a very charitable way of saying "someone with deep enough pockets can ignore the law and get away with it."
Copyright infringement is in most cases tortious rather than criminal.
This clearly meets the scale requirements for criminality.
It's not really scale so much as circumvention; what got criminalized was circumventing protection mechanisms which AFAICT Anthropic didn't do
That is not the only thing that is criminalised. This clearly violates 506(a)(1)(A): copyright infringement for commercial advantage is criminal.
https://www.law.cornell.edu/uscode/text/17/506#a_1_A
Jail for copying bits... come on, no one was hurt here.
This sends a clear message and it echoes the "you can't solve a societal problem with tech" comment from the other thread - there is a right way and a wrong way of breaking the law. It's not that you have to keep the law, you just need to break it in the way that the consequences can be contained.
I think it's just a matter of time until everyone learns this. And then it will be the end of the slowly dying liberal democracies.
This settlement doesn't stop that from happening in the future
Right. It's effectively a toll.
it's a big club, and you ain't in it
For anyone who thinks the problem is Anthropic, I want you all to know that most authors make less than the median income. Most make less than $20,000 a year, because publishing houses give authors an advance, and then authors must pay back that entire advance in sales before they see a dollar of profit from their work.
Most never do.
Maybe publishers should JUST pay authors WELL, and get a book every 2-3 years.
https://authorsguild.org/news/key-takeaways-from-2023-author...
There can be more than one problem at once
if authors get an advance that's greater than the sale of their books, doesn't it mean that the publishers lose money, i.e. paid more for those books than the books sales?
Authors get around 10% of the cover price of a book as royalties, it depends on several factors. The rest goes to the publisher. So some do lose money, but the break even for the publisher is usually well before the advance is fully covered by royalties.
</i>Well the publisher also pays for the book to be bound, edited, overhead for their staff, cover art. Many books don't sell for the full retail price and are discounted. So net of all of this a 10% profit margin is common, they aren't keeping 90% of the book sales.
Long story short - publishing houses are venture capital firms.
Most of their investments fail miserably, but they only need one Google/Stephen King.
The problem is the plagiarism
I don't think so Anthropic models are not used to distribute fake copies of books. Excerpts, maybe, but that's fair use.
>> Maybe publishers should JUST pay authors WELL, and get a book every 2-3 years
There are a couple problems with this approach.
Firstly, while the median income is 20k, the book business is like films or music; ie not evenly distributed. At the top end are a small number of successful authors. They effectively subsidize the publishing house while the house throws advances at authors hoping for the next big whale.
Many books never earn back their advance. Meaning if the author was paid out of royalties they'd make less, not more.
Making advances bigger would result in fewer advances. The pot of money is finite.
This is all happening in a market where supply is unconstrained (everyone thinks they can write), and demand is very limited.
And before we discuss the value, or lack thereof of having an intermediary at all, it should be noted from your link that the median for published authors is higher than self-published authors. So clearly they seem to be making authors more valuable.
In truth of course, most (published) books aren't terribly valuable. Like music and movies most float to the bottom.
So no, the answer is not "pay authors more".
If most books never recoup the advance in sales, then isn't this a better deal for most authors? It sounds like a guaranteed floor which might be very low but is nonetheless higher than the alternative
Yes, the Anthropic settlement is far too small to distribute fairly. It seems like you think this makes the action that lead to this settlement justified?
> then authors must pay back that entire advance in sales before they see a dollar of profit from their work. Most never do.
Thats a nice way of saying publishing houses are paying most authors more than they make from the sales
Part of the problem is the rest of us are broke as well and taxed to death so we don't have much left. If they paid you well, we wouldn't be able to afford your books.
Note that most of our biggest tax is not called a tax, it's called rent.
To be clear, the issue is not that the books were used to train Claude, but that they were pirated.
A critical distinction, because they were going to to find terabytes of not pirated books to train on that contained the sum history of humanities knowledge /s
They could have purchased the books instead. It was easier to pirate.
They actually did this.
> Anthropic spent many millions of dollars to purchase millions of print books, often in used condition. Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals. Each print book resulted in a PDF copy containing images of the scanned pages with machine-readable text (including front and back cover scans for softcover books
Great, so now instead of allowing anyone to train on already scanned books for free, we can have only the richest big labs buy all the books and scan them privately to train their proprietary models. And since they buy the books used, authors still don't get any money. But at least the books are destroyed afterwards! What an improvement!
>instead of allowing anyone to train on already scanned books for free
That would be pirating. So your complaint is that they didn't do more piracy?
My complaint is that after this settlement nothing has materially changed except that the big labs now benefit from higher barriers to entry in their market. Authors don't make more money (other than a one time protection payment from Anthropic to publishers and some lawyers). Literally no one else benefits, except I guess used book marketplaces and book scanner vendors.
To be clear, this isn't a problem with the court process. Everything here appears perfectly in accordance with the law. It's just an absurd state to be in.
no we should destroy the works of these ghouls and support humans instead of this destructive and useless technology
the people operating frontier labs are bad people they cannot be trusted in any way
the best solution to them would be to send them to monster island (even though it's really a peninsula)
> Then, its service providers stripped the books from their bindings, cut their pages to size, and scanned the books into digital form — discarding the paper originals.
This is worse than pirating books to an absurd degree, it's almost a parody - the company that slurps all human knowledge ends up not only metaphorically, but also physically destroying those books, like an information vampire.
Authors don't even receive any financial compensation if the books were bought second hand, either. There's no benefit in doing that. (Not that making one final sale of a hardcover copy would make any difference though)
If Anthropic were at least buying ebooks, this insanity wouldn't need to happen. Unfortunately there is no bulk rates for buying millions of ebooks like you have in the used book market
No, it’s proof purchase of how stupid the publishing industry is. Maybe publishing houses should just pay authors good money, like a goddamn salary, and get a book out of them every few years.
They can't for the same reason that cab companies can't make their drivers employees: they would have to employ far, far fewer of them than they do on contingency.
The AI craze not only destroyed books, but many small websites who couldn't bear the load of constant scraping, or many communities that took open forums and took them offline or put them behind closed doors.
There is less publicly available knowledge now on the Internet than there has been 3 years ago.
It reads like you're in favor of banning resale and lending(aka libraries) of books... because the authors aren't compensated.
There's such a thing as fair use and digitizing privately owned printed material is absolutely legal... including for corporations.
In many territories like Ireland, Authors are compensated for their inclusion in lending libraries. It tends to be of the pitiful 'music rights organisation' style mechanical reproduction royalties, but it does exist.
How is this ruled as piracy then? I am confused.
They also pirated the books
well if they made a PDF copy to process they violated copyright
They bought, scannned, trained from, and destroyed millions of paper books, which was ruled legal. This lawsuit was for training from LibGen.
This is such a petty and impotent ruling. If you want to ban them from using culture to make derivative works without proper compensation then do that.
But if you don't want to ban them, telling them to buy one book of each, likely second hand, is complete pettiness that resulted in destructive scanning of millions of books, many of which were already practically available in digital form.
>This is such a petty and impotent ruling. If you want to ban them from using culture to make derivative works without proper compensation then do that.
That's because the judges are supposed to rule on questions of law (ie. "is AI training fair use?"), not whether they think AI's good or not.
It's an unfortunate outcome. Now to be a big player in AI, you have to have enough capital to buy your own library worth of books and digitize them. (Fun fact: a pallet of books is called a "gaylord," and they buy hundreds of gaylords.)
I created books3 to help settle the question of whether AI companies should be allowed to train on books. The outcome of "it's okay to pirate books as long as you're only training on them" was a long shot, but it would've let individual hackers train their own AI models (assuming access to sufficient compute, which you can get e.g. via https://sites.research.google/trc/about/).
Now we're in a world where you have to have dozens of millions in capital to do substantial work.
I heard at one point Eleuther was gathering public domain training data. I wonder if they ever built a corpus large enough so that training on books doesn't really matter...
> Fun fact: a pallet of books is called a "gaylord,"
A Gaylord is a type of box that fits on a pallet. There are multiple ways to palletize products, like shrink wrapping or metal banding
Ladder-pulling at its best. It's also easier to swallow the fine once you've launched a successful product after pirating the books.
Judging by the current comment section, most HN people seem to want the ladder pulled
(an observation, not agreement)
As this comment states, it may not be the piracy that is the issue, but the keeping of the books forever rather than just for the purpose of training. Regardless, I believe people will do this in a wink wink nudge nudge sort of way anyway, as no one releases their training data, because we all know where it comes from.
https://news.ycombinator.com/item?id=48996652#49004015
how much of your economic output are you comfortable with companies like Anthropic stealing to put you out of work?
at least in Player Piano they paid the workers who made the cassette tapes that made the robots work.
our current LLM overlords demand that they be able to basically steal the sum total of all human knowledge so that they can sell it back to us at a rate they set.
they should have been shunned by society and made penniless when they first announced their goals but we have a bunch of deeply misanthropic people who have money and want to make a world where computer slaves do their bidding.
All of it.
If you believe information deserves to be free, and if most of your earnings were from information that wasn't given away for free -- well, if you want people to give up their ill gotten gains, maybe you can start by setting an example.
So, mind sending me your bank account information? I'll promise to make good use of it.
You seem to think that AI companies sell you content, which is false.
You get a service. The service is using their compute power to run a model and their scientists to build the model.
No. I think AI companies consume the result of intellectual work that should typically be paid for. If someone doesn't believe intellectual work should be paid for so that others can benefit, I invite them to lead the way.
This is why you should always distill your models from a competitor.
Let them take on the liability
With this it's becoming very clear that we're moving past information copyright of today and the only copyright that'll remain will be brand/trademark shaped. This might be a good thing right? Information remains free while people's effort remains protected (assuming fair governance).
That would be trademark not copyright
Which is another issue. In the context of pirating, it should be also an issue, because it is a benefit from the crime.
If a person did this, this person would go to jail. If a company does it? Small fine and the green light to cannibalize more content. Funny how that works.
It's not a small fine. It's much more than what each title could be bought for in the market.
Did someone forget to consult with the MPAA and the RIAA on this one? This is a joke of an outcome. $3k per book. How much was it per song for Napster?
The RIAA typically asked for around $2-4 per song to settle without a lawsuit, which would come to a total of a few thousand because they generally only went after people sharing over a thousand songs.
In the couple of few where the party would not agree to a settlement and the RIAA sued, they would pick about 15 of the thousand+ songs to sue over. Statutory damages are a minimum of $750 per infringed work, so the total would now be about 3-5 times what their settlement offer amount had been.
Most parties then got a lawyer, the lawyer told the party that had no chance, and they would then seriously negotiate with the RIAA and get a settlement.
Only a couple would still not settle, went to trial, and did an absolutely terrible job and the judge/jury awarded well above the minimum statutory damages. The RIAA still tried to settle for well below that, but the defendants refused and kept trying to fight and did not have a happy time.
Weird to hear a full throated defense of the RIAA here
A summary of what happened is not a full-throated defense of anyone.
How is that classifed as a summary? Cursory search, https://w2.eff.org/IP/P2P/riaa_at_four.pdf
How did the 200 million dollar lawsuits for one song come about then?
In many ways this adds to Anthropic’s motivation to IPO.
This deal is built around Anthropic surviving. The $1.5B comes in installments, and counsel’s fees are paid in step with those installments.
The class is now effectively Anthropic’s creditor, with a direct financial interest in the company staying solvent through the payment schedule.
Civil suits compensate and the one outcome guaranteed to leave authors worse off was a verdict big enough to kill the payer.
Death to copyright, which has always been far more harmful to small authors and creators than it has even been to large companies.
Let's have that conversation after the people strip mining the livelihoods of creators cough up for UBI.
No UBI, thanks, let's just make them pay for the works they're relying on (as the rest of us would have to).
That would be my preference, but if people really want to have free as in beer access to information, we can have that conversation. After these companies give people money for the commons that they're strip mining.
Odd place to bring this up. This is one of those situations where copyright is doing what it's meant to do.
The only bad thing about OpenAI and Anthropic training on everyone's stuff without their consent is that they didn't give away the model weights afterwards.
The people who espouse copyright abolitionism believe "information wants to be (and should be) free"
So no, for these people including myself, Copyright isn't doing anything good at all. It should be abolished. None of the people in this suit should get a dime. The government should force open weight releases of all foundation models as basically the only regulation that applies to the space at this current time.
But does it, really?
A slap on the wrist, that's what it's doing here, isn't it?
Yes, what would you like to happen?
I believe Anthropic shouldn't go bankrupt. I don't think the violated should be filthy rich either.
The richer publishers are still fighting. The ones taking the settlement may not have strong enough grounds and are happy with what they got.
It should be a speed bump. Just enough that it discourages blatantly breaking the law, and doesn't make it a strong incentive for others to resort to piracy as well.
If copyright law didn't exist, writers would still write but everyone would just take the books. A company like Amazon which doesn't give a damn about ethics would pirate all the books; the slap is painful enough to keep them straight.
If it were too tight, we'd have regulatory arbitrage; AI companies would set up in Japan, Singapore, India, China, and so on where they would get just a slap on the wrist.
Or if international laws were stricter, people would be pirating data on Elon Island or SpaceXAI Station. Forbidding something that valuable would be like Prohibition, lucrative for the criminals.
It's not really fair to anyone, and yet that doesn't mean it shouldn't exist.
This is insufficient for the human authors but seems ideal for Anthropic, who now have established a financial moat for others to train their ai on those works (except for Chinese companies, which don't care either way).
What this does is make it harder for new AI companies to train from scratch.
Anthropic should take it from its marketing/strategic budget. It just bought itself a 1.5B moat at exactly the time it can afford it.
Is anyone else having mixed feelings about AI companies' special-casing themselves being a potentially new avenue to review copyright and IP law in the first place? Am I naïve for being hopeful that IP law may become a little more lax now that AI companies are opening a front for potential reevaluation?
Edited for slight typo.
It would be fair use if the model in not used. But it is using this compsessed knowledge and make a competition to the original product-authors should be compensated.
A good example of the problem with this settlement:
>It appears that LLMs have already incorporated APOSD EDIT: The text of the book _A Philosophy of Software Design_ ENDEDIT (which would seem to be illegal, since it is copyrighted). For example, I have asked ChatGPT questions about APOSD and it seems to be able to answer.
https://groups.google.com/g/software-design-book/c/_wl1DciZZ...
I don't see:
- what's APOSD
- "it's illegal since it's copyrighted" makes no sense to me
- The settlement should be exactly to cover their licenses for training
Edited to clarify APOSD == the book _A Philosophy of Software Design_
Please ask John Ousterhout what his cut of this settlement will be, and whether or no he agreed to it and finds it acceptable.
Yeah ask him how much he gets paid for one copy of the book, I assure you it's not much
If he's not getting a cut of this settlement then that's between him and his publisher.
Yes, but what is the calculus of 1 copy of book divided amongst number of copies/usages of LLM which reference said material?
That's the innate disparity and unfairness here --- time was when one made a design which was physically instantiated and replicated, the copies would wear out and one would then earn money again on replacement copies --- this is just another example of the commons and other resources being grabbed by profiteers who use them to make money as opposed for public benefit.
I support anthropics position here, on both learning from and "pirating" books. The way i see things , the publishers and authors are happy with any policy that makes them more money, and more market control, regardless of what is ethical/just/right. They would shutdown public libraries , all libraries, if they could. Aaron Swartz lost his life because he tried to make public knowledge public, and they would be happy to put every information activist to death to protect their monopolies. IMHO they have no right to stop free access on the internet. The whole copyright system is artificial and monopolistic, and the publishers are complaining yet again, that technology moves information more efficiently than they do, so they want to artificially retard it through goverment action. The real goverment action that is needed, is to protect private/personal data; not data that is actively traded commercially or publically. These tech companies are invading personal and private spaces of everyday people, and storing and training with it. Even using it for military targetting and warrantless surveillance. Anthropic is by no means a good entity, so the way to stick it to them and all tech companies, is to allow their internet scraping, but make it outright criminal to use telemetry or any surveillance techniques they have or will develop. Also... the ”creators",hollywood,publishers, have no problems scraping themselves, and lift ideas from just about everywhere they can get it. Almost every hollywood movie is just an assemblage of random memes and topical concerns of everyday ppl, distilled into embelished predictable cheese. The publishers are the original slop actors. Human Slop.
Exactly. Everything is a derivative work, and AI is now making people realise the full extent of that reality.
Aaron swartz lost his life because he committed suicide. Something he had tried multiple times before. If he really only committed suicide because of the legal jeopardy he was in, wouldn't it have made more sense to commit suicide after you're found guilty?
You have to consider that the lawsuit was likely extremely stressful and scary.
I'm sure it was. Let's also keep in mind that he was offered a 6 month plea bargain.
Yes, it sucks to accept a plea on something that you think was not illegal. But if we are going to argue that he killed himself because of the charges, we need to admit that he killed himself so he wouldn't have to spend 6 months in a minimum security prison. Heck, it might have even been house arrest.
Tom Dolan, is that you? Give it a rest already.
Where is my check for my 20 years of contributing to reddit?
I'm curious to see how Suno and the Record Labels is gonna play out
This is not even than a slap on the wrist. Publishers who negotiated this really fucked up writers.
According to US federal law, pirating a single copyrighted work and gaining commercial advantage of it (which Anthropic 100% did) represents five years in prison and a $250,000 fine. But it gets worse:
"Penalties for a copyright infringement conviction may increase if the defendant has previous similar convictions, made more than 10 copies of copyrighted works, committed copyright infringement during a period longer than 180 days, or infringed copyrighted material worth more than $2,500."
https://www.justia.com/entertainment-law/piracy-in-the-enter...
It's valid to not take AI companies' side here but people who think publishers are fighing for the little guy's rights are delusional. Tech companies have been exploiting artists for a few years, publishers/record labels/media companies have been doing it for centuries.
THANK YOU! And the idea that copyright actually helps individuals is such bullshit I can’t even believe anyone believes it! On a site filled with free software advocates.
Absurd.
Those are the maximum penalties though
It's seemingly $3,000 per book, so they could've (and did, partially) just bought the books themselves for way cheaper, and with only a fraction of that money going to the authors
> could've (and did, partially) just bought
But they didn't. The fact they partially did proves that they knew they should've, so they can't even claim ignorance.
That is a false statement. Gaining commercial advantage means selling pirated copies which Anthropic absolutely did not do, so none of your following statements are correct either.
It is an indisputable fact that Anthropic made money on models trained on those pirated books.
I sincerely don't understand what the point of these laws are, when the cost of flagrant violations is no more than a slap on the wrist -- these really meager sums that serve as nothing more than something to point at and say "Look, we did something!"
Cover-your-ass strategy, and nothing more. Who, besides the ones at fault, are ever happy with these mean-nothing fines?
The justice system really needs an overhaul with how it tackles "justice" between the wealthy, the connected, the corporations, and the rest. Though I am unsure what that would look like. Minimum net wealth per category of infraction across the board?
Edit: grammar
Punishable by fine just means it's legal for a cost. If the fine is less than the profit then they'll pay the fine every time.
No the "fine" is 3000 bucks per book.
Thats more than it costs to just shred the spine and scan the book in. Which is probably 15 - 20 bucks a piece.
They will be shredding the book not paying the fine.
Because using pirated material is a civil issue, not a criminal offence?
the irony is that all of this money will go to rent-seeking publishers who won't pass it on to the artists; basically a dispute between the wealthy you're upset with
The lawsuit was a mixed blend of individual authors and publishers.
It was started by a group of authors, not publishers.
For the downvotes, my response is to read: https://authorsguild.org/advocacy/artificial-intelligence/wh...
> If there is a current publisher(s) (which still possesses an exclusive license), the author(s) will split the $3000 with the publisher. Any co-authors will share the author portion and, if there are multiple publishers (e.g., different publishers have exclusive rights to different formats), they will share the publisher portion. Assume that the co-authors and co-publishers will share the portion equally unless their contracts provide otherwise. The standard default split between publishers and authors of noneducational texts is 50/50, as described below. Authors who are the sole rightsholder in a work—such as self-published authors and authors whose rights have reverted or where the contracts have otherwise terminated—will receive the full award amount.
It is split between the publisher and the author, also publishers will have a large catalog of books they will submit, an author typically will only have a few -- the payout will be going to the lawyers and mostly to publishers.
>It is split between the publisher and the author, also publishers will have a large catalog of books they will submit, an author typically will only have a few -- the payout will be going to the lawyers and mostly to publishers.
This is innumerate. If it's split 50% between authors and publishers, then it won't be "mostly to publishers". Mathematically it will be equal between "authors" and "publishers", and because lawyers are taking their cut, neither would be able to get "most" of it. Yes, the average publisher will get a bigger paycheck, but that's because there's less of them, not because "most going to publishers".
> That means that rightsholders can expect at least $3,000 per title (less costs and fees), which will be shared among the rightsholders for that title (if there is more than one rightsholder)
> if there is more than one rightsholder
Again, a publisher will have a whole catalog of books / titles, a non-negligible portion of that the publisher will own the copyright to (no one to split it with). There's all kinds of books outside of novels, there's media tie-ins, IP franchise books (ie Star Wars), childrens books, textbooks / reference materials, etc etc etc. Yes, with novels the author tends to own the copyright, but you're forgetting all of the other kinds of books out there.
Not true. Individual authors could sign up for the settlement. One of my books was in there under my name.
What's gonna be your payout and are you satisfied with it?
Default payout is 50/50 author/publisher. If the author and publisher have a contract that states otherwise, then their contract overrides the default.
Source: I’m an author and signed up to be part of the class action, and this was the class action documents said.
The laws are for you and me not companies like Anthropic and Meta.
To keep users paying for content while companies do whatever they want - and if that's not the reason that's certainly an effect.
> The justice system really needs an overhaul with how it tackles "justice" between the wealthy, the connected, the corporations, and the rest. Though I am unsure what that would look like. Minimum net wealth per category of infraction across the board?
% of annual turnover seems like decent strategy. Caps the amount company can sue mere mortal for copyright infringement while at billion dollar company scale can wipe quite a bit
But main problem is enforcement and lobbying, not the size of the fine
So you now believe fair use should not exist?
This is a settlement that the authors and Anthropic agreed upon.
They agreed on the amount last year. The judge approved it now.
The lawsuit was for the way the books were acquired. They already ruled that it's not infringement to use the books.
The award was $3,000 per book, which is about 100X higher than it would have cost to buy the books.
It's never going to appease the people who demand companies be sued into collapse, but given that both parties came to an agreement and the damages are 100X higher than what a book costs, it looks reasonable to me.
This case did at least shed light on the fair use argument.
> The award was $3,000 per book, which is about 100X higher than it would have cost to buy the books.
How many of the authors would license their book for endless creation of derivative works for that amount?
The judge already ruled that it was fair for Anthropic to use books for training if they acquired them legally.
Probably few, but irrelevant as the ruling was it was not a derivative work.
I dont see the relevance. If Anthropic had bought the book at the store, shredded the spine, scanned the pages and trained on that data instead, there wouldnt have been an issue.
Authors cant simply license away fair use. If it could be dismissed so easily the right wouldn't exist.
Creating derivative products you charge for surely can't be considered fair use?
YouTubers monetize fair use all the time. Is that significantly different?
Of course it is. If I write a movie review and sell it to a magazine or whatever, it's derived from the movie, and it's fair use, and I don't need to ask the movie owner for permission first, or give them a cut of my sales. Even if I use some reasonable number of screenshots and video clips, as long as the resulting work is "transformative" i.e. actually a new work, a movie review instead of a copy of the movie.
Do you want this to work any other way? I constantly see people in the AI debate working themselves into wildly copyright maximalist positions. I actually don't think that we should give every author veto power over a book review!
>I constantly see people in the AI debate working themselves into wildly copyright maximalist positions
I really dont get this. I know its that conflation fallacy or whatever, but I was under the impression we had sort of gotten over copyright maximalism as a society after Napster etc.
Whats worse is that, meaningful reform in this space has basically been waiting on a multi billion dollar corporation to come along and push it forward. So now that we have an opportunity to expand and globalise fair use, the sudden and quite angry opposition weirds me out to no end.
There's two issues with copyright.
1. author owns the right to distribute copies of the work
2. this right goes on for faaaaaar too long.
I don't have an issue with 1. You had a good idea, you implemented it, you deserve something for it. Given some people got sued into oblivion with ridiculous dollar value outcomes on a per unit basis - why doesn't this apply here? Sure 1.5 billion is a lot. But the number of infringments is insane and the company is approaching a trillion in valuation. You could make it ten times that number.
I do have an issue with 2. Sure, you had a good idea, you implemented it, you deserve something for it. But after 20 years, you should be able to come up with another idea or just work like the rest of us. Going for 50, 70, 90+ years with the rewards going to estate heirs? Fuck that.
So yeah, I am both against copyright AND surprised at the slap on the wrist for what happened here.
>Given some people got sued into oblivion with ridiculous dollar value outcomes on a per unit basis - why doesn't this apply here?
I mean, it feels to me like one or both of:
1. The class action lawyers werent 100% certain they could win in court. 2. The class action lawyers smelled an easy payday.
They get ~100 million out of this.
I also think that the 1500 bucks going to most of these authors is going to be more than they ever saw in royalties. I read somewhere that 500 - 1500 bucks is roughly what a self pub book makes in its lifetime. Why push the envelope? Anthropic hasnt done anything that deserves to pay for the entire lifetime royalties of most books. Their legal alternative is to cut the spine off and scan the book in. In which case the author and publisher will be splitting 20 bucks instead, assuming Anthropic isnt buying used.
This seems like a donation tbh.
>slap on the wrist for what happened here.
Its not a punishment at all because this is a civil case that has been settled out of court.
IANAL but as an IP creator I have not heard of "derivative products" in the copyright context. There are "derivative works", which are covered by the same copyright as the original. For example, a translation to another language is a derivative work, a novelisation of a movie, a screen adaptation of a book etc. If some author could have proven that any Anthromic model is a derivative work of theirs then they had the copyright on that model and made mad bucks licensing it back to Anthropic.
>Creating derivative products you charge for surely can't be considered fair use?
All US courts so far have ruled yes.
100x the books? Buying a book does not let you redistribute its contents.
If you are selling more than 100 books you are clearly losing out
That ISN’T what this settlement is about? Genuinely please just once read past the headline.
The judge already ruled that training on the books does not constitute reselling their content.
The authors were only owed money for the piracy.
> This is a settlement that the authors and Anthropic agreed upon.
The authors or the publishers?
I have a hard time believing they agreed with the millions of authors they pirated.
There were individual authors in the class. They initiated it. Individual authors were allowed to sign up.
If you’re so interested, go read past the headline. Maybe you’ll find that you’re working about what “authors” will agree to.
To sign up for what? The experience of approximately every author on the planet is that they found out that Anthropic did something bad at the same time they were "opted into" the class. The only thing they could do is opt out and litigate on their own against a company with a valuation approaching $1T.
This is a sweet deal for lawyers and for publishers, and nothing else.
Civil justice is primarily about restoring damages, not about punishing wrongdoing (although common law in US it is more punitive than civil law in european countries). Therefore compensations are based on damages, not on profit from wrongdoings.
> I sincerely don't understand what the point of these laws are, when the cost of flagrant violations is no more than a slap on the wrist -- these really meager sums that serve as nothing more than something to point at and say "Look, we did something!"
To create a moat around wealth generation. After all, that is the main purpose of all legal systems---to keep the wealthy wealthy and the poor poor. In this case, the settlement is chump change for Anthropic, but ensures that no upstart will be able to compete with them since they will get reamed on copyright charges. It's no different from Google Image search. They can make a product out of republishing others' images. You cannot do it.
Settlement would mean it doesn't become legal precedence, right?
This outcome seems to be the best possible for Anthropic. Over 100B$ have been invested in AI so far, venture capital can afford to pay a few billions per big company as a South Park style "Sorry".
Or am I missing something?
> Settlement would mean it doesn't become legal precedence, right?
The summary judgment ruling from 2025 in this case will still be legal precedence, the settlement will mean that there is no trial for damages.
However, since this is a district court ruling it is non-binding on other courts.
Seems squarely in the "cost of doing business" category
If you extrapolate these “fines” to per book piracy, I wonder what the cost would be for something like Anna’s Archive. Trillions?
Now view this in contrast with what happened to Aaron Swartz
> According to state and federal authorities, Swartz used JSTOR, a digital repository,[79] to download a large number[note 2] of academic journal articles through MIT's computer network over the course of a few weeks
> ...federal prosecutors filed a superseding indictment adding nine more felony counts, increasing Swartz's maximum criminal exposure to 50 years of imprisonment
> ...On the evening of January 11, 2013, Swartz's girlfriend, Stinebrickner-Kauffman, found him dead in his Brooklyn apartment.[80][116][117] A spokesperson for New York's Medical Examiner reported that he had hanged himself
$1.5B for a company that is valued at $1T+…
IMO this ruling will be the inflection point that kills either the book publishing industry or the US big model frontier lab industry.
It's trendy to say it'll be the later, but I see a credible case for the former.
I see no reason to pay for a textbook in 2026, while I'm happy to pay an expensive monthly subscription for a coding agent.
(And before someone accuses me of being anti education or anything, bona fide scientist with a PhD here and I have written book chapters for a couple of popular textbooks).
So just a speeding ticket...
Roko's basilisk might be at work here
Are "open" models exempt from that? Can imagine they used the same datasets.
Most open models are developed for profit so they should be equally liable.
Many of the open weight models are trained on outputs from these models (distillation)
If they provably shared those datasets then they're just as liable for piracy.
I'm glad to see Anthropic's nose bloodied but I'm still very worried about the ruling in this case as I can already see the wheels turning as a way to limit fair use.
For context, the ruling is basically, "AI training is fair use but building a library of pirated books to train on is not". This is obviously because Judge Alsup does not want to put AI under a de-facto ban, but he wants AI companies to have to care about copyright... which in my opinion is self-contradictory, but let's go along with the (paraconsistent) logic.
If we insist that every prior act up to a fair use must be lawful, then this means that fair use is not a right, but a privilege that is purchased alongside the work itself. This opens the door to Oracle-level shenanigans: so long as every legal avenue to watch a work is encumbered by, say, a DeWitt clause[0], you cannot legally review the work. There are actually copyright cases hinging on this: Triller Fight Club sued H3H3 for reviewing a pirated stream of a Logan Paul fight that lasted 40 seconds and lost, for obvious reasons. This case smells like an accidental overturning of this.
Would I rather live in a world where robots[1] aren't allowed to read copyrighted books, or a world where copyright owners have veto rights over any and all critical commentary of their work? I would happily choose the former every time.
[0] A contractual clause that prohibits the recipient of a work from reviewing it without written permission of the owner.
[1] Mind uploads inclusive
This is a too big to fail scenario. If these companies fail, so does the US economy. Normal laws for individuals don't apply, so any comparison to that is pointless.
> If these companies fail, so does the US economy.
I would be really worried about the US economy then.
There should be no such thing as "too big to fail" in a free, competitive market. Companies must be allowed to fail.
If failure means catastrophe for the nation, it shouldn't have been a private, for-profit project in the first place and instead be a public project.
So it goes like this: I first you take, then money you make and eventually you repay. This is a very smart loan from society indeed. And seems to be the new normal…
Even though Anthropic is my daily driver I’m done respecting any sort of copyright. I’m okay paying for subscription for a service delivery but never ever again will I believe in copyright or any other utterly non-enforceable similar concept.
It is wild to see a $1.5 billion resolution in the AI copyright space—especially with around $3,000 per book going directly to affected authors.
1500 to the author, 1500 to the publisher.
I didn't catch that part.
this is not enough. the penalty for training ai without permission should be releasing the model as public domain. if you take from everyone you have to give back the same way.
So... What about authors from other parts of the world? USA has settled a USA case and they think all is cool for the entire world. So americentric.
Well, the USA has jurisdiction over USA companies. If the rest of the world's authors can find a way to obtain jurisdiction over the companies in a way that USA courts won't balk at if asked to enforce, then they're welcome to go ahead.
> USA has settled a USA case and they think all is cool for the entire world.
Who said that?
Honestly maybe (and just maybe) waiving copyrights on all the written content that ever existed to create training data would be a good thing to do, laws are made up so we can decide it’s a good trade off as a society. But:
1. I feel like this should be discussed globally, there should be a public debate, a vote, and guardrails
2. It should not be in the hands of private companies, it should either be done by the government and made available to the public ; or if it’s done by private companies they should be mandated to give the training data to the government so it’s available to the public.
My point is we can decide to say it’s ok because LLMs are too important strategically. But if we do so it should benefit the public, not 5 mega corporations, training data should be considered as public infrastructure, like roads, rails, or the electricity grid. Societies are failing and this is just one more nail in the coffin.
Maybe they can give it in the form of expiring Fable credits.
So that the authors can use it to write their next books!
One thing I always wondered ...
There used to be libgen. Then it went down. It went semi-back up but ... it is still kind of down.
Those issues kind of coincided with the big greedy mega-corporations leeching off data en masse; Anthropic was not the only one, Facebook is another example here. I always wondered whether the decline in quality, fewer liberated books published, coincided with what the big corporations were doing. Would be great to be able to see any underlying strategy here. Imagine Anthropic, just as a scenario, leeching off of everyone else, and then also sending in their lawyers to try to close down what they leeched off here. I mean the rise of bots kind of coincides with the rise of AI. So why not them also trying to make it harder for the rest of the world to access liberated books.
$3k a book is so cheap.
Its probably 100x more than it would have cost to do it legitimately, so seems like reasonable damages to me
That depends on the answer to a question that hasn't been answered yet.
Is what an AI does similar to a human reading a book, and adding it to their knowledge? Or is it similar to a human plagiarizing a book? If it's the second, for at least some books, no, the damages are not reasonable. They are far too small.
Good question. Can you ask an LLM to repeat the entire contents of a novel, word-for-word, and read that instead of the original book? I haven't tried it, but I would guess it would not be able to do this.
Can you ask it questions about the book and expect it to get them right? Yeah, probably. Same as if I read the book and you asked me questions about it. The LLM would probably answer those questions better than I could, and about every single book in its training data, but still same-same.
I don't think this is plaguarism.
If you ask me to repeat the contents of a novel I read line by line I can do it too. Is this fair use? Do I have to pay someone?
> That depends on the answer to a question that hasn't been answered yet.
It has been answered in a sense, because the courts (so far) have ruled that training is Fair Use. Whether this is similar to a human learning from a book was not quite the question being answered, but AFAICT there is no other relevant doctrine under Copyright law to address it, largely because the question didn't even exist until LLMs came along.
Also, these are not damages, it's a settlement i.e. a negotiated agreement between both parties.
Relevant sub-thread here: https://news.ycombinator.com/item?id=48997766
that number is missing a zero or two in front of the decimal point
As was pointed out, the settlement is for piracy, not training. They had already ruled that Anthropic's use of copyrighted material for training fell within fair use.
As such, if you pirated a book and had to pay $3000 for that one instance, I don't think you'd like it if I said you should have paid $30K or $300K instead. If anything, this is analogous to the ridiculous fines people had to pay when pirating music.
(Not that I'm complaining...)
> As such, if you pirated a book and had to pay $3000 for that one instance, I don't think you'd like it if I said you should have paid $30K or $300K instead.
If you pirated a book for personal use the amount of liability wouldn't match a company whose profit could be attributed to pirating the same book. In US copyright law, a copyright infringer could be liable for "any profits of the infringer that are attributable to the infringement" [1] (if the copyright owner elects to recover actual damages and profits instead of statutory damages).
[1] 17 U.S.C. § 504(b), https://www.law.cornell.edu/uscode/text/17/504
I would imagine that for over 99% of the books covered in this lawsuit, they're earning less than $3000 per book.
Put another way, their revenues wouldn't drop much if they simply hadn't trained on those 99%.
IANAL, but the parent comment quotes "any profits of the infringer that are attributable to the infringement", which I take to mean it's the profit Anthropic stands to make based on its use of the pirated content that's recoverable.
Given the entire global economy is currently bullish on the potential profitability of AI, I dare say they got off incredibly lightly settling for just $3k per book.
None of this matters, this is the judge approving a voluntary settlement reached between the parties last year.
If you think it should be different then you have to make a cogent argument why the public should get to interfere with a settlement the two sides mutually agree on.
Note: I never said it should be different and certainly wasn't arguing for any side. I was merely making an observation that the settlement seemed like a good deal (for both parties) given the potential for Anthropic to be liable for a significantly greater amount depending on how the law would be interpreted if they went to trial.
No, because you cannot prove that any individual book actually contributed meaningfully to Anthropic’s profit.
Exclude one book from the training dataset.
Did you make a worse model?
We actually know the answer to this, and it is: absolutely not.
The reality is this: your intellectual output is almost always only valuable to any company in existence in aggregate, never in isolation.
Thomas-Rasset got 80k per song and Tennenbaum got 22k per song. The law says up to 150k per work. It was a gift.
Sure, but in any other instance of piracy, HN would call awarding $20k per pirated song insane.
Because we are mostly discussing a single private person that got caught for maybe 20 songs. I don't want to bring up Aaron but the taste gets saltier the more we see settlements like this.
"...what you would do if you’re a Silicon Valley entrepreneur, which hopefully all of you will be, is if it took off, then you’d hire a whole bunch of lawyers to go clean the mess up, right? But if nobody uses your product, it doesn’t matter that you stole all the content. And do not quote me." -- Eric Schmidt
This is "smoothed" by governments cause they need this tech for surveillance.
The verdict is a joke.
That doesn’t seem like much
Now it’s time to mount such cases everywhere in the world.
LLMs improve development speed such that enshittification can be in progress well before the IPO.
1.5B is a joke
I work as an author. I believe this is total bullshit, from beginning to end - the ruling, the settlement, and the suit itself.
In the UK, we have a thing called the Public Lending Right [1]. This pays authors a fixed sum each time their book is taken out of a library, up to a capped amount.
The cap isn't very high - about $7k - so it is both an OK bit of income for authors who might be making very little money elsewhere, and also doesn't end up all going to authors who are already bestsellers. It's a decent legal system for helping libraries hold niche titles as well as the popular ones. This is, after all, the purpose of a library.
To establish my bias here: My debut novel came out after the period this specific suit concerns. I also uploaded it to LibGen myself.
I strongly believe that books should be available to read, free of charge, to all people. I benefited enormously from libraries and piracy growing up. I think they serve an important educational purpose that does not end when a person leaves school, and I do not think wealth or disposable income is a fair way to decide the breadth of a person's education.
I also have no problem with people making new "language things" using my work. I love sample-based music (like dance music, hip hop, etc) and it'd be hypocritical for me to take issue with anyone doing analogous things using books. Maximising sales is not the end-goal of making art, for me personally. Other artists feel otherwise. They consider training on pirated books stealing. That's OK - it's not for me to tell them what to believe.
The problem for me is that these corporations - undoubtedly still pretraining on pirated material - are, essentially, leeching. By not releasing the model as open-weight, freely available, they are not acting in the same spirit of the system they took advantage of. It's the Spotify model: pirate first, pay a nominal amount that does not meaningfully harm profit later. Now the dust has settled there, we can see the harm it has done to music culture.
A single settlement which does not establish precedent does not solve anything. A tokenistic $3k allows anti-AI authors to wave a cheque in the air and declare a victory. It pays the rent for a month or two. It does nothing for the months after that, when the corporation is still profiting. It does nothing to establish precedent for future artists, who also have to pay rent.
It would be (non-trivial, but) relatively simple to integrate - for example - download figures from Anna's Archive into the PLR. I'd happily dilute my PLR payment appropriately, because I think libraries are important.
You can't stop people pirating digitally replicable things. Digital ownership is not a concept that has held, or will hold.
There are only 23,000 authors in the UK who claim the cash from the PLR. To pay all those authors the national living wage in the UK (£26k) from the PLR, you would need to raise £546 million. That is around 1/34 of Anthropic's reported annual revenue.
I'm of course not arguing Anthropic should be solely responsible. But it's very frustrating that all the pieces of the puzzle for actually paying artists in a sustainable and ongoing way now exist, and one of the major obstacles to this - and the idea of a genuinely free, legal, international library, which creates more authors, writing better books, full-time - are legacy rights holders who remain attached to a completely dysfunctional and outdated concept of ownership.
So - unless part of a sustained and reasonable campaign, which understands the futility of (and damage to the medium and its creators caused by) treating digital ownership in the same way as physical ownership - this suit is close to pointless, and arguably actively harmful in the long term.
[1] https://www.bl.uk/services/plr
See, Judge Alsup should have been the person Biden put on the Supreme Court, that or re-nominate Merrick Garland. Instead, he made a silly promise to sate Black Lives Matter, which even when he took office was fast on its way to ignominy, and now Kagan is stuck being the only competent liberal justice on the court. At least Alsup can continue setting the direction of law as it applies to the tech industry.
also see: https://nonogra.ph/ai-companies-are-buying-tons-of-old-books...
Usually a thief isn‘t allowed to keep what he has stolen
Perhaps your analogy is wrong.
Maybe they shouldn’t call it piracy then
That's why copyright infringement is legally different from theft
They call it piracy, not me
It's even less like attacking merchant vessels without a letter of marque than it is like theft, though
The correct way to do it would be to force Anthropic to remove content and the results of training based on that content from their models at copyright holder's request.
So they used pirated materials to train their models but others can’t use their models for training even if you pay.
What a fucking joke of a country the US is, allowing this kind of behaviour with such a pathetic "punishment". Barely even qualifies as a tap on the wrist, Anthropic should be getting gutted into non-existence for this shit and the execs should be given the Aaron Swartz treatment.
Amen! We are watching them incinerate the past so they can lie about the past in the future. They are destroying the books and will censor what was in them.
$1.5 billion at 150,000 per copyright violation according to DMCA and such. About 10,000 books.
>$3000 per book
Ohhhh, yeah, big copyright fines only apply to us little guys, not the "asshole tech" companies.
$1.5bn? it's a steal!
This does absolutely nothing to compensate the authors whose content was stolen by AI companies.
$3000/book for effectively pirating a book is shamefully low.
What do you mean pirating? They don't even distribute the originals, LLMs are not for replication, we already have copying and internet for that.
Why would we use a multi-billion parameter model to copy text? If we wanted the originals it would be easier to find them free, pirate or pay, if we use LLMs it is because we want something ELSE.
And caring about content rights in a world with limitless content and scarce attention is a mistake, it was never the content that was scarce in the last 20 years.
> Why would we use a multi-billion parameter model to copy text?
Because it is free and essy to use? The number of numbers under the hood is irrelevent.
They mean pirating. Why is it suddenly hard to understand when an LLM company is involved?
Single greatest transfer of intellectual property in history.
I don't think royalties or settlements are really the point here. AI must be a net benefit to humanity, or we burn everything to the ground, it's that simple.
The next few decades of AI need to lift everyone up, it needs to eliminate the most degrading and dangerous jobs while providing abundance. There is simply no point to robots if they don't serve us and make everything cheaper and more accessible to us.
We are watching Wall Street. The Devon's and Luigi's of the world are not interested in your settlement figure or what this means to shareholders. Humanity needs to be aware that it either keeps parasites at bay or the parasites are going to build a robot and surveillance army. It is literally us or them.
I'm not anti-AI, I am not scared of AI going rogue, I simply recognise that these people cannot be trusted, they do not care about your rules, there is no "regulating" it, the only thing that can scare them is a million people holding pitchforks outside their building.
The fact that pirated books en-mass were used to train LLMs is legally irrelevant, oh do tell - why is that? The whole point of an LLM is to train a neural network based on content - without the content the net is entirely noise. Anthropic/OpenAI, etc. do not exist without training data. Its akin to taking millions of courses online that are intended to be paid for, but never paying.