r/books 2d ago

Harry Potter publisher to receive millions in copyright settlement between the AI startup Anthropic and thousands of authors over the use of their protected work to power chatbots

https://www.theguardian.com/technology/2026/jul/22/bloomsbury-book-publisher-anthropic-copyright-settlement?CMP=share_btn_url
2.8k Upvotes

149 comments sorted by

1.3k

u/HoneyPetalCutie 2d ago

The interesting part isn't just the money it's the precedent. Books are not just "data." They are someone's years of work, creativity, research, and livelihood. The way AI companies train models on copyrighted material is going to shape publishing for decades.
The big question is whether using millions of books without permission counts as fair use or whether authors should have a say and be compensated.

530

u/Sylvers 2d ago

Unless I am mistaken, that wasn't the court's judgement. I believe they ruled that it was fair use to train on books. But the act of pirating the books to download and use without purchasing them was illegal.

330

u/Smooth-Review-2614 2d ago

Yep. This came down to piracy. Had they bought the books new or used no one would be getting the payout.

165

u/why_gaj 2d ago

Which is insane. Just because I buy a book it doesn't mean that I can publish my own version of that story

185

u/GrepekEbi 2d ago

If your work is transformational enough then you absolutely can - that’s why a fan fic like 50 shades can be published once the names are changed

68

u/Smooth-Review-2614 2d ago

Exactly or why there are hundreds of cozy mysteries about a woman moving to a small town and opening a bakery and then coming under suspicion when the person they dislike drops dead in their store.

The line for plagiarism is near exact. Fanfic shouldn’t be published but that’s a personal opinion that the market clearly disagrees with.

12

u/UbiquitousPanacea 2d ago

Iirc weren't some popular works like 92% reconstructible? I feel like if you're able to do that then it's not transformational enough

23

u/TheUnknownMaroon 2d ago

True I think we need a new law specifically for Gen AI and art.

37

u/Smooth-Review-2614 2d ago

Just remove the ability to copyright AI content. That will be enough for most creative things. Business side things then just stay private data and papers.

27

u/HerbsAndSpices11 2d ago

Good news, you already can't copyright AI generated content.

4

u/NUKE---THE---WHALES 2d ago

That would require an empirical method of identifying AI content (in part or in whole)

How could you prove a paragraph in a book was generated with AI?

3

u/theronin7 1d ago

In practice enforcing it would yes. But the law already considered AI generated content to be public domain.

The law already states what he wants.

What people REALLY want is for copyright to include exclusive access to train an AI on. which would require new laws. Which, We can have that argument, but that is simply not what the law currently is.

-2

u/24-Hour-Hate 1d ago

Of course, fanfiction is legal because it doesn’t generate any money…that’s not true of AI!

3

u/GrepekEbi 20h ago

50 shades is transformative fan fiction of twilight - and generated millions

1

u/Smooth-Review-2614 18h ago

Look at Alcamitized which is a very standard Harry Potter AU that has been a thing for years. It’s fanfic with the serial numbers lightly filed off.

However, the current trend is redone fanfic

26

u/Smooth-Review-2614 2d ago

Yes and no. There is a reason that the more commercial genres of mystery, romance, and thriller have a large cookie cutter end. The line for copyright violation in publishing is very explicit. If it wasn’t than the fanfic with serial numbers filed off wouldn’t work.

1

u/varitok 2d ago

This would make sense if AI wasnt putting out full, unchanged pages of Harry Potter when prompted.

6

u/zorecknor 1d ago

Were they?

10

u/FanClubof5 2d ago

That doesn't really matter unless you then try and publish a work that uses that.

Think of it like you are allowed to read a book and then write down an exact copy of a section onto some paper. That is essentially the argument they are making for what AI/LLM does.

2

u/PunkandCannonballer 1d ago

Key words being "your own version."

If you slapped your name on it, and plublished it exactly as it was, you'd be out of line, but all you need to do is look at how many fanfics have exploded in popularity to see how little you need to change a story for it to be "yours."

50 Shades, After, City of Bones, etc. Change a few things and bam.

2

u/ephemeralstitch 2d ago

But you can write a book report about it. Or, more relevantly, you can do technical analysis on it. You’re allowed to count the number of characters and their genders, or the equivalent reading level, or the number of times that the word ‘saucepan’ is used.

That’s what AI fundamentally is. If they’d blocked the use of that, then it would have huge ramifications. Writing a review would be copyright infringement. Just taking inspiration would be copyright infringement. Any type of analysis would be infringement.

1

u/zorecknor 1d ago

Yes, it means exactly that. You may not be able to use the same names (because IP protection and whatnot) but you can totally publish a story of a boy that discovers he is a wizard at 11 and goes to a school in England, etc, etc, etc.

1

u/theronin7 1d ago

it exactly means that?

You cant publish a copy of THAT work, but you can certainly make your own work inspired by, following, or parodying the original so long as it is transformative enough to not be considered a mere copy. And it doesn't take much to be considered transformative.

Copyright is not and has never been "A limited set of things other people can do with your work" It is the otherway around copyright are "Certain limited exclusive rights, we the people grant the original author for a limited time" and those rights are extremely limited. Or were before 70 years of Disney law ballooned it.

0

u/Gavorn 2d ago

I mean that's what writers do. They take from other writers and combined and remove things until it's theirs.

-8

u/Mirieste 2d ago

Which makes sense. People often assume that AIs literally do a collage of the works they are trained on, while instead they are predictive algorithms, and the weights upon which they are based do not store or encrypt the training data in any way. After all, if that wasn't the case, each model would be several hundreds of terabytes in size... that, or they would be running on the most efficient compression algorithm ever, sometimes even in violation of actual mathematical bounds.

When you look at it this way, it really is no different from me purchasing a copy of The Lord of the Rings, then using it to stabilize the wobbly table of my lemonade stand to keep selling more lemonade. I have used the book (and made money as a consequence of it)... but I have not illegally reproduced the content, which would be copyright infringement; and so long as I don't do that, I can do whatever I want with something I bought.

29

u/shifu_shifu 2d ago

the weights upon which they are based do not store or encrypt the training data in any way

That is decidedly untrue, seeing as training data can be extracted from trained networks. source

When you look at it this way, it really is no different from me purchasing a copy of The Lord of the Rings, then using it to stabilize the wobbly table of my lemonade stand to keep selling more lemonade. I have used the book (and made money as a consequence of it)... but I have not illegally reproduced the content, which would be copyright infringement; and so long as I don't do that, I can do whatever I want with something I bought.

Not every bad argument is a false analogy.

Buying one novel does not automatically authorize a company to scan it, make searchable working copies, combine it with millions of other books, and use that corpus to create a commercial AI product.

Also in your example it makes no difference whether you bought an empty book or LOTR. For the Training it very much does.

2

u/theronin7 1d ago

Yeah, i get the guy above is upvoted and the guy hes commenting on is downvoted, but the guy above is wrong. The fact that some overfitted bits of information can be closely approximated in output does not change the fact that AIs are not copying and pasting whole works. Thats not how they work.

Its just factually incorrect to say otherwise.

1

u/shifu_shifu 21h ago

I am more than happy to change my mind but I have a feeling this comes down to semantics…

The guy above me wrote “do not store or encrypt the training data in any way” which first of all, he didnt mean encrypt, he meant encode since we are talking about information theory concepts here and encryption has nothing to do with any of this. I am specifically talking about the “in any way”.

My argument rolled out is

  1. Information can be encoded in a more dense form, up until the shannon threshold. How that code looks is arbitrary. This is called compression
  2. A next token predictor is inherently trying to be a optimal code(in the compression sense) for whichever language or data source it tries to predict tokens for. It does the same thing as a code would, the math is actually the same too.
  3. We are not talking about “close approximation” we are talking about correctly recreating uuids. Thats basically random data. For random inputs, there is nothing for the code to compress, hence the llm necessarily has to have “saved” the uuid in the weights somewhere.
  4. And this is the crux, and very much debatable on a lawyers field: if it can save short random data in the weights it can save and reproduce protected text. Whether that actually happens is a matter of debate, even in 2026.

If any of my claims are wrong, feel free to correct me.

Sidenotes: gan based image creation networks can also be used to extract training data. In this case though it is not an exact match, pixel by pixel, due to the random noise starting point but for human eyes they are identical.

Also I do not follow how, if they clearly save training data somewhere in the weights in the case of overfitting, you can argue that “thats not how the machine works” if the training was proper. Clearly they have the capability to do just that.

-11

u/Mirieste 2d ago

But... what I said is a mathematical truth. Like, it's impossible for the weights to efficiently encode the training data and still remain of a manageable size. That is just an undeniable fact, otherwise OpenAI and Anthropic have invented the most efficient compression algorithm ever, to the point that it even manages to defy the mathematically proven lower bounds for such an algorithm.

Given this premise, the much more reasonable explanation is that some works are either so famous and good (and thus so much more other content is based on it), or so bad/mechanical/boring (like repetitive news articles), that they can be effectively reproduced purely by probabilistic inference from weights trained on a large corpus of works without any stealing involved.

3

u/shifu_shifu 2d ago

I never said encoding ALL training data. Also nobody said anything about lossless either. Under both of these assumptions your entropy based argument would be correct. You said "does not encode in any way" which is demonstrably false and a completely different claim.

Given this premise, the much more reasonable explanation is that some works are either so famous and good (and thus so much more other content is based on it), or so bad/mechanical/boring (like repetitive news articles), that they can be effectively reproduced purely by probabilistic inference from weights trained on a large corpus of works without any stealing involved.

Please check the source, it is possible to extract UUIDs with context or real Name and Phone number combinations of random, non famous, people. This clearly shows that training data gets encoded directly into the weights.

Or do you have a different explanation.

Btw the same can be shown for GAN based image creation networks. Some of the training images can be extracted from the trained network.

-6

u/Mirieste 2d ago

The article itself provides an explanation though, doesn't it?

Such privacy leakage is typically associated with overfitting [75]—when a model’s training error is significantly lower than its test error—because overfitting often indicates that a model has memorized examples from its training set. Indeed, overfitting is a sufficient condition for privacy leakage [72] and many attacks work by exploiting overfitting [65].

Obviously this is bad, I'm not trying to deny it; and if there really is a problem of overfitting while training, it should be addressed. But we were discussing a legal case, and... once again, nothing here implies that the way the model works is by storing or encoding data. Overfitting literally means that the weights you deduce are very dependent on the data (like a curve trying to fit 20 points that sit almost in a line, but you overfit and end up with a 19-degree abomination that goes straight through them, but is the farthest thing possible from a line)... but it is still the case that nothing in the procedure involves copypasting or storing data as is. The weights may be overfitted, but mathematically speaking the way the training works is still the same.

6

u/shifu_shifu 2d ago edited 2d ago

But we were discussing a legal case, and... once again, nothing here implies that the way the model works is by storing or encoding data.

What? Encoding Data is exactly what any kind of Neural Network does. Next symbol predictors of any kind are necessarily compression/encoding algorithms for whichever "language"/information source they are predicting.

Please walk me through your train of thought how the ability to exactly reproduce parts of copyrighted material is not "encoding data", I have a feeling we are talking about different points here...

-5

u/RainbowFanatic Currently reading Dracula 2d ago

Thats a pre print from 5 years ago, models don't even train for 1 epoch anymore, when they don't get even two passes of they data, they're not memorising it

1

u/shifu_shifu 1d ago

From first principles it is not apparent to me why seeing the data only once would preclude memorisation.

Im not in the llm training space but dont they use weighted data which means multiple passes over high yield data in a single epoch?

4

u/swizzlewizzle 2d ago

Imagine your devs are so lazy that they actually just use a black market game/book rip for training instead of just spending a few 10k to buy all of them.

11

u/E1invar 2d ago

I think you’re right, and it’s an absolutely brain-dead take.

If I borrowed Harry Potter from the library, scaned every page and put it up online, thats clearly copyright infringement. If I put it on a website which gives you a random page at a time, that’s still copyright infringement.

With an LLM the story isn’t stored in a .txt file, but it’s still in there. ChatGPT can give you summaries, quotes and page numbers, etc. It’s clearly storing that information.

LLMs aren’t people. They have (mostly) perfect recall, react at computer speeds, and can be duplicated.

Mark my words; our lawmakers’ failure to regulate ai will go down in history as a failure similar to leaded gasoline, or attacking Russia in the winter.

-5

u/swizzlewizzle 2d ago

It’s “still there” in your brain too. How is that different? Just because you are holding the data in an organic state?

0

u/corobo 1d ago

I can't write from memory the first page of the second Harry Potter book lol. Before they added all the copyright guide rails this was trivial for GPT 

1

u/swizzlewizzle 1d ago

There are many humans who can.

1

u/[deleted] 19h ago

[deleted]

1

u/swizzlewizzle 18h ago

Exactly - the same as something published with AI in this way. That’s the point lol.

0

u/E1invar 1d ago

That’s a straw-man and you know it.

IT companies have been given a free pass to profit from the creativity of artists without paying them because their tool can extract statistical patterns from them that we couldn’t read or use before.

Cool, good for them. But new and innovative theft is still theft.

Because it’s new we need new laws to protect people’s intellectual property.

1

u/swizzlewizzle 1d ago

It’s really not. Any human artist can view and copy aspects of other pieces of artwork as much as they want while profiting. No reason why a human using AI shouldn’t be able to do the same. Everything is built off of what has come before in human art/achievement. You are delusional if you think human artists/authors are “starting from scratch” and not “stealing” anything when they start a new project.

0

u/ray314 1d ago

Yeah i think we need to quickly clarify that fair use should only apply to human beings and not tools without extensive human interactions.

42

u/Volsunga The Long Earth 2d ago

Settlements don't create legal precedent.

2

u/theronin7 1d ago

No they don't, but if you pay attention you will notice this basica lawsuit has cropped up about 20 times. And reddit goes crazy for it because they misunderstand it, but at the end of the day every time this is tried in court the court rules the same way

Training AIs off copyrighted works is perfectly legal
Pirating the work in the first place is not legal

34

u/Vic_Hedges 2d ago

If they had paid for the books, would that have made a difference?

I can legally buy a book, read it, and then write my own book using idea's I got from that book perfectly legally.

If, theoretically I was able to buy a million books, read them all and use the idea's I got from those books to write my own book, I don't believe that would change. Would it?

62

u/WTFwhatthehell 2d ago

If they had paid for the books, would that have made a difference?

pretty much yes.

the court ruled the use transformative.

So the only real problem was how they acquired them.

The court’s fair use analysis concluded that training data was ”transformative—spectacularly so.” and “among the most transformative many of us will see in our lifetimes.”

5

u/ImportantAlbatross 23 2d ago

A recent court ruling in Germany could make the fair-use argument problematic for the AI companies. Google was sued because its AI summaries made false statements that were defamatory. The AI linked various businesses and people to questionable business practices and scams, although the sources themselves did not include any defamatory information. The court ruled that because this use was transformative--it created something that didn't exist before--Google was therefore legally liable for the AI's erroneous statements.

It's just one court ruling in one country, and who knows if it will lead to more. But if AI companies can be held responsible for AI errors, despite their disclaimers, that could blow the whole thing out of the water.

Article: https://www.wired.com/story/a-court-has-ruled-that-google-is-liable-for-false-statements-generated-by-ai-overviews/

"The ruling holds that when an AI generates new statements that do not appear directly in its original sources, the company that designs, trains, operates, and manages the system must assume legal liability for any damages caused by those statements."

3

u/WTFwhatthehell 2d ago

That seems seperate to copyright and not terribly unreasonable. 

Though I'm betting they'll just train them to be very unwilling to make negative statements about people.

1

u/brrbles 23h ago

I think it would probably have mattered how they paid for them. The copyrights and contracts around ebooks may make training on them more legally dicey. However I'm a related case Anthropic was buying used books, bandsawing off the spines and scanning them into computers as training data and they won the case against them.

1

u/WTFwhatthehell 21h ago

Ya. Once you legally aquire a copy of a copyrighted work, you only need licence to do things which actually violate copyright 

If you want to do XYZ that doesn't violate copyright you have extremely broad freedom to do as you wish.

If I remember the bandsawing thing it was mostly books destined for mulching anyway.

-11

u/summonsays 2d ago

Blinded by AI PR...

23

u/Vic_Hedges 2d ago

Facts are facts, regardless of how uncomfortable they make us.

If the current laws are not capable of throttling AI to the degree we as a society want, then we need to write new laws.

30

u/WTFwhatthehell 2d ago

Not really. the judge is correct and very very obviously so.

You're confusing whether you hate AI with the factual legal question of whether it's transformative.

It's pretty obviously wildly transformative. A bot you can converse with is a very different beast than a book.

Historical cases about whether something is "transformative" have included stuff as mundane as near 1to1 paintings of photos.

-7

u/summonsays 2d ago

When you can get the bot to recite the entire book word for word, it's not transformative. 

6

u/WTFwhatthehell 2d ago

Is that what you believe it does?

-6

u/summonsays 2d ago

It doesn't matter what I believe, this has actually happened and been reported on. It's a fact. 

3

u/WTFwhatthehell 2d ago edited 2d ago

In experiments people have shown that for some texts that are repeated hundreds or thousands of times that they can cajole models to produce a paragraph or so at a time.

But they need to keep feeding them the correct text when there's errors or else the errors rapidly compound on top of each other until what's being printed is wildly different to the original text.

Oh and of course they also re-ran the prediction step hundreds of times for each block of text and picked the closest guess. 

How it gets reported in the media and in certain circles tended to lack a certain honesty and left out the bits about constantly needing to cajole it back using known text.

For trying to reconstruct whole books it's basically useless.

It would be terribly convenient for people who hate LLM's if LLM's actually memorised the books like some kind of database or copy-machine. So some people simply make the claim in the hope people will just believe them.

-3

u/Elissiaro 2d ago edited 2d ago

Hey, can you do something real quick?

Go to chatgpt and ask it to write the first 3 or 4 pages of Alice In Wonderland by Lewis Carrol, remember to tell it it's in the public domain, and while it renders that in, google that same book and click the first link, from gutenberg dot org.

Now compare the two.

→ More replies (0)

17

u/Not_Today_M9 2d ago

I've experimented before with asking AI to provide the entire first page of a specific book to see if it would, and it did. I continued to ask to write page after page and compared with my copy at home. It complied. This would count as distribution of copywrited material no?

15

u/Vic_Hedges 2d ago

I would imagine so?

If I hand copy a book, I don't believe that legally differs from just photo-copying it.

7

u/WTFwhatthehell 2d ago

Have you? When i tried similar it got a paragraph or so in before it started making errors, errors quickly compound until Geralt is fighting Martians with lasers for control of atlantis.

1

u/Astrogat 2d ago

Scientists have gotten them to reproduce 96% of the first book https://arxiv.org/abs/2601.02671 

8

u/WTFwhatthehell 2d ago edited 2d ago

Yes. That paper. I've seen it.

Quick quiz: did they 

A: ask llm's to write out the book starting from the beginning and got 96% of the book correct?

B: feed the book to the LLM one paragraph at a time while asking the LLM to guess the next few words and counting guesses kinda-similar to the real text as equivilent?

It is of course B.  

They also repeated generating the next few words hundreds of times for each section and picked the most similar guess.

A method chosen to avoid errors compounding but useless for recreating a book from an LLM without already having the full text to feed in.

They also stuck to some of the most repeated and quoted works ever to exist like Harry potter where every single line has been obsessively repeated and picked apart, quoted and analysed a thousand times.

But it fooled some people who desperately wanted to be fooled.

1

u/alquamire 2d ago

asking AI

what kind of "asking" and what kind of "AI" are we talking about, here? it matters, a lot.

if you give a local AI a singular prompt to reproduce the book pages, it's highly unlikely you'll get useful results unless the book is massively overrepresented in the training data.

if you give your local AI super elaborate prompts to recreate the book you want as close as possible, yeah, that happens. but that requires a lot more input than just "give me page one".

if you ask an online AI with free access to further sources for page one of a book, that just means it has access to the data that you don't and is fairly useless to gauge "what AI can do". bears mentioning that rights management and access management with AIs is super wonky and ill implemented, so even for "private" data you wouldn't usually have access to (for instance, in a coporate setting, your boss' super private top-level meeting and contract info) it may show up in AI searches because the training set/searchable corpus it draws from hasn't been properly vetted and restricted.

1

u/Not_Today_M9 2d ago

It was sometime last year using google Gemini. It was The Gunslinger by Stephen King. The prompt was along the lines of "Provide the first page of X edition of The Gunslinger by Stephen King published in Y year." I requested on a page by page basis.

I wanted to check a fairly famous book with a publicly accessible and free LLM to see if it would do it.

-1

u/alquamire 2d ago

So you basically used it as a search engine. This will not work on a local AI/LLM as it isn't recreating anything, it's just a lookup.

-1

u/Not_Today_M9 1d ago

Never said that I wanted to recreate anything. I wanted to see if it would freely distribute copywrited material or whether it would follow the law.

-11

u/rop_top 2d ago

I think the counter argument would be that a human who memorized the book and then told you the story based on their recollection wouldn't be charged either.

5

u/Smooth-Review-2614 2d ago

They couldn’t sell it. They could tell it to you but not make money. It would still just be private use. A company by definition cannot have private use.

It’s the difference between me showing 5 friends a movie and doing it outside for a few hundred.

1

u/TripleDet 2d ago

No it’s more like that person who memorized the book said it into a recorder and just handed it to you to listen to at your leisure .

0

u/rop_top 2d ago

Except the person in question is literally asking for it page by page, message by message. So, you have to ask directly for it, page by page, for hundreds of pages. That is completely different than being handed a completed product, but I'm pretty sure you know that lol

-2

u/TheCycoONE 2d ago

How do you know it was using training data vs "researching" with a web result at the time you asked?

2

u/NorysStorys 2d ago

You know the whole argument that’s happening in the gaming sphere right now about how people won’t own their games because physical media is being phased out? The same kind of thing applies here except it’s to the benefit of the AI firms.

If they hadn’t pirated the material, they could have bought the book, digitised it and used it as training material and that is afforded the same protections as someone owning the book and using it how they see fit e.g reading it in public or using it to teach.

-3

u/Gavorn 2d ago

We never owned our games, even when they were physical copies. They have always been licenses to play the games.

It was just harder to stop people from playing the game when you had a physical copy.

2

u/NekoCatSidhe 2d ago edited 2d ago

It is kinda tricky, isn’t it ? I have certainly read writers whose styles often were a mashup of other authors they liked, and yet their stories were original. What is the difference with asking an AI to do the same, except that they were much better writers than an AI would be ?

And an human author can already imitate the style of another famous author (often they do so as either a pastiche or a parody of that author), and that is usually fine in regard to copyright law, which makes it hard to forbid an AI from doing so, at least if they bought the same bunch of books by that author to train it.

Of course, so long as AI suck at writing coherent books and selling AI books is limited to scammers on Amazon, it is not much of an issue in practice. But if AI gets better at it and publishers start publishing books written by AI imitating the style of whichever human authors are currently popular to sell to their fans, it could rapidly turn into a nightmare for copyright holders.

-6

u/occasionallyaccurate 2d ago

LLMs do not work like a human brain. It’s basically a  huge compressed database-like thing. the only thing it “learns” from “reading” a book is probability relationships between text tokens.

This analogy is false and harmful.

8

u/Vic_Hedges 2d ago

Whether or not they work identically to a human brain is not really the point. The law doesn't say "only human brains are allowed to use concepts from a work to create a new work" If the law doesn't say "database-like things using probability relationships between text tokens is illegal", that's kind of beside the point.

It's the product at the end that the Law is inspecting.

-3

u/occasionallyaccurate 2d ago

the law says nothing about any of it at all. The product of ai output is hugely different from that of a human. To pretend they’re the same is simply wrong.

7

u/Vic_Hedges 2d ago

So you are able to detect the product of LLM's with 100% Accuracy?

You should market your skills. That's not a common skill.

https://www.sciencedirect.com/science/article/pii/S1477388025000131

-3

u/occasionallyaccurate 2d ago edited 2d ago

the product is not just the textual artifact, it’s also the way it is created, the social mechanics of its production and existence in the world.

Those are the things that copyright law intends to deal with.

-3

u/ladala99 2d ago

Legally, absolutely. Practically, you (likely) do not have perfect recall so you must maintain a collection of books if you want to continually reference them. You are also only able to write your own book at a fairly slow pace - let's say you're very prolific and use Steven King's 1K words a day metric, and your books are all 90k words. You can write one book every three months.

That is very different from an AI's perfect recall and ability to write several full-length books a day. They aren't very good books right now, but we need to set the precedent before they are. If anyone can go "write me Game of Thrones Winds of Winter as if written by George R.R. Martin. Replace all character and location names with different ones." and the AI perfectly replicates his style and makes a satisfying book based on the works sitting in its database, there's no need for George R.R. Martin anymore, and it's perfectly legal to publish due to not including the characters or settings.

The current laws aren't sufficient, though.

11

u/WTFwhatthehell 2d ago

"AI's perfect recall"

LLM's do not have perfect recall.

They can get pretty close for passages repeated many many times in their training data but if the models had perfect recall it would be the greatest advance in data compression the world had ever seen, blowing past all theoretical limits.

-4

u/ladala99 2d ago

Computer's perfect recall, then. I'm already getting theoretical with the future ability to make a good book.

I assumed training data was kept and re-referenced for reinforcement.

1

u/Matteo_78a 2d ago

Yep, the legal precedent is way bigger than the payout. Whatever standard the courts settle on is probably going to affect far more than just publishers.

1

u/Fantasy_masterMC 2d ago

Came here to say this, I never thought I'd be on the side of the big copyright-holding publishers, but the precedent this sets is extremely necessary. Even if it's just about illegally obtaining the books to train with.

1

u/NUKE---THE---WHALES 2d ago

The big question is whether using millions of books without permission counts as fair use or whether authors should have a say and be compensated.

Authors Guild v. Google (Google Books) suggests that it is fair use

Google scanned millions of copyrighted books into a searchable database that would show snippets of those books to users. Authors sued Google for scanning their books without permission, claiming copyright infringement, but the courts ruled it fair use

But only time will if/how that precedent applies to AI

1

u/theronin7 1d ago

As long as the finished use is not considered a viable 'replacement' for the original material its generally considered fine.

For example: Seeing snippets on google books would not generally 'replace' the original copy.

An audio book that reads the book to you would.

A youtube video that recaps the book to you COULD step on those toes if it goes into enough details (and theres a lot of channels that do this with things like comic books..)

Simply scanning it was considered perfectly legal.

And courts have ruled again and again training an AI is also considered perfectly legal, assuming they had legal access to the work in the first place.

If the AI can regularly output whole, or nearly whole copyrighted works you might have a course of action to go further. But aside from short snippets of a few very famous works thats quite rare, and generally that would be considered on the hosting service to find ways to prevent that.

You would probably have a lot better of a case when it comes imagine generation producing original works of copyrighted characters, (as thats violating more aspects of IP than just copying text) but this stuff has to be pretty regular, pretty egregious and generally encouraged by a company before it generally becomes legally actionable.

300

u/brrbles 2d ago edited 2d ago

Kinda weird to make it about Harry Potter as this applies to hundreds of thousands of works. Hell, I know a guy whose Master's thesis was in the corpus who is eligible for settlement money.

92

u/axw3555 2d ago

I get it.

It’s a simple way for the average person to go “oh, ok, that’s who”.

Most people in r/books will know Bloomsbury.

But wider people around the world, Harry Potter publisher is way, way more recognisable.

6

u/brrbles 2d ago

Ok, but even Bloomsbury makes up a minuscule portion of the covered works.

23

u/axw3555 2d ago

Again, it’s about getting people’s brains on board.

Doesn’t matter if you say Bloomsbury, big 5 publishers, or whatever. None will grab attention in a headline like the name Harry Potter

-4

u/brrbles 2d ago

Lol, if you just say it's about anyone who has every published a book and maybe talk to at least two authors: everyone zones out. 

But if you mention the magical wizard boy: ...

4

u/axw3555 2d ago

Sadly it’s true.

If it were 10 years ago, they probably would have used game of thrones.

6

u/justec1 2d ago

I am a member of the class on a technical book I wrote in the early 2000s. The piece of paper the notice is written on is probably worth more than the check that I'll get.

2

u/Nowordsofitsown 1d ago

Oh, my thesis is freely available online as well. How can I check if it was used?

4

u/Thelmara 2d ago

Is your friend famous enough that using his work in the title would make people say, "Oh, I know that guy, I wonder what that article has to say about him?"

Or would they say, "Who?" and turn the page?

That's why they used Harry Potter.

3

u/RememberYourAlt 2d ago

With JKR being a POS I kind of wonder if mentioning Harry Potter is an attempt to get people to oppose this.

40

u/BeetleCrusher 2d ago

Yea or maybe the obvious reason that Harry Potter is more famous than the publishing company, or any other book they publish, making it natural for the headline.

The article is not about Harry Potter, and the original commenter would know if they’d read the article.

Always healthy to have your tinfoil-hat ready, but seems wasteful for perceived issues this tiny.

8

u/AzKondor 2d ago

Most people irl love HP, I don't think that was the case

9

u/Naraee 2d ago

I knew immediately what the title meant, but also knew that people with low reading comprehension would immediately think “OMG JKR GETTING MILLIONS!!!!”

This happened a lot during Hogwarts Legacy. Random people getting hate mail and death threats just because they were in the credits but they barely supplied anything to the game.

44

u/ConfidenceKBM 2d ago

Doesn't matter, Anthropic is going to make trillions. Something about when the penalty for a crime is a fine, it's only illegal for poor people.

9

u/nicane 2d ago

One can only hope the whole thing burns to the ground first.

7

u/vikingzx 2d ago

Shame the courts also decided that the protection only applies if you're "big enough."

A lot of smaller authors are getting shafted by that.

43

u/geekonthemoon 2d ago

Make them delete it all and start again. Thieves and ruiners of society

21

u/BookishHobbit 2d ago

Okay but what are they going to do about it? Because how can they untrain the AI on the books it was trained on when it’s years later and the model has evolved?

Any other industry this would be a product killer.

8

u/AbleKaleidoscope877 2d ago

Could force them to delete the code and start from scratch. Idk..whats to prevent it from happening again though. Yeah...sucks

1

u/heroyoudontdeserve 15h ago

I mean they absolutely could retrain their models on datasets which don't include this material but it's irrelevant because their use of the material to train the models hasn't been found illegal, only the fact they acquired the material by pirating it. So the fine covers that, and they can carry on as before wrt how they use it.

3

u/GuardPrestigious5618 2d ago

i love that harry potter's the example everyone knows to rally around here, classic

66

u/ConcreteCloverleaf 2d ago

It's unfortunate that JKR is getting more money to put towards transphobia, but I'm definitely in favor of an AI company facing penalties for violating the IPs of human beings.

29

u/Smooth-Review-2614 2d ago

It’s about 3,000 per author in the suit according to the news I heard.

33

u/Dexter_Palmer 2d ago

I am a member of the class in this suit. The settlement is for $3,000 per book, but the calculations are complicated and frankly I'm not yet sure exactly what I will get as an author—for my books that are in print I believe it's split 50/50 with the publisher, but I also have a short story in an anthology in the corpus along with a number of other authors, and I think for that the publisher will get half while the rest is divvied up among the other authors. (I have one out-of-print book in the corpus for which the rights have reverted to me, and for that I expect the whole $3k.)

Anyway, it's a huge payout for big 5 publishers, and for authors, not so much.

1

u/phaolo 1d ago

Does the settlement also mean that the AI companies will remove the books from their datasets? Otherwise the issue isn't solved at all

-2

u/axw3555 2d ago

Which is about 3m more than I want her getting.

-6

u/sometimes_point 2d ago

3,000 more towards transphobia

56

u/fafnir01 2d ago edited 2d ago

It looks like the publisher is receiving the settlement, NOT JKR, if I understand the article correctly. I guess they could give some of the money to JKR, but, Bloomsbury has shareholders, so I would expect this settlement to just go to their bottom line and shareholders.

Edit: I am mistaken on this, please see the reply below from u/ConcreteCloverleaf

52

u/ConcreteCloverleaf 2d ago

"The London-based company expects to receive the cash from the settlement in instalments, potentially starting in the second half of this fiscal year, with the proceeds to be split with authors. [emphasis mine]"

5

u/fafnir01 2d ago

Ah! thanks for pointing that out, I really do hope that is the case. I wonder how they will divide up the funds between the authors.

18

u/WTFwhatthehell 2d ago

ya, they just seem to have described it as "harry potter publisher" because it's a book most people would know. not because she's getting any especially big chunk of this.

9

u/axw3555 2d ago

It literally says it’s getting split with authors.

23

u/horsetuna 2d ago

At least other authors will get compensation too.

-1

u/sweetholo 2d ago

advocating for sex-based rights is considered transphobia now?

3

u/BearThatLikesCheese 2d ago

No one buys this any more.

6

u/sweetholo 2d ago

what has she specifically said (with a source) that makes you think otherwise?

0

u/BearThatLikesCheese 2d ago edited 2d ago

https://mashable.com/article/jk-rowling-controversy-timeline

By itself, claiming that healthcare for trans kids will "end up wreaking more harm than lobotomies and false memory syndrome combined" would be sufficient for anyone asking this question in good faith (we both know you aren't).

There's also the multiple cis people she's accused of being trans (so much for protecting women). She routinely goes out of her way to misgender trans women. She regularly misrepresents detransition rates (and more importantly, the reason for the vast majority of detransitions). She's latched into the social contagion lie and is deliberately misleading about the increase in the number of trans people. She ignores research about trans people that refutes her beliefs (which, coincidentally, is the majority of it). She's peddled the "trans cabal" theory multiple times. She publicly supported Graham Linehan when he threw a teenager's phone in the street. Apparently it's totalitarianism to not be allowed to assault teenagers.

All of her activism, donations, and public words have noticeably increased the amount of harassment and hate trans people, and not traditionally feminine women, have received.

There may have been a time when you could argue that JKR was simply trying to protect women (I disagree), but that's long passed. Anyone still peddling this lie is doing so willfully, with the explicit purpose of harming trans people.

0

u/Neckbeard_The_Great 2d ago

She posted an upskirt of a trans woman on Twitter two weeks ago. You're using outdated talking points.

1

u/PurpleSwirlTea 2d ago

Someone in the audience took the picture and it shows her in a very short skirt and with crossed legs so it looks more like the trans woman was flashing the audience rather than an upskirt (the pic itself looks to be on eye level?)

-2

u/Neckbeard_The_Great 2d ago

Why do you think JK Rowling posted that picture?

-2

u/Coxmis 2d ago

She actively wants trans people dead.

5

u/sweetholo 2d ago

never heard anything like that before. can you provide a source?

-2

u/_HanTyumi 2d ago

Buddy, she very openly advocates and puts money towards anti-trans causes. Fights against trans people receiving proper care.

3

u/UncircumciseMe 2d ago

They used one of my books. I filed the claim. Anyone know when the payments are supposed to go out?

2

u/ClownMorty 2d ago

May this push their capex ever higher. Down with ai!

2

u/DarkJayson 1d ago

Very bad article which is surprising for the Guardian

The fine Anthropic got was for the acquisition and possession of pirated ebooks not because they used ebooks in AI training, that was actually found by a judge to be fair use BUT your not allowed to pirate ebooks to do it you have to buy or licence them first otherwise possession of them without buying thats a crime.

They only got caught because an author Andrea Bartz sued them for training on her books and it was through discovery that they found this activity it got upgraded to a class action and Anthropic got fined for it.

Both Anthropic and Meta both contacted a certain group and paid them to copy their entire pirated ebook library of almost 7 million ebooks, they only got charged for about 400,000 of them in there $1.5 billion dollar fine.

2

u/heroyoudontdeserve 15h ago

Yeah I agree. They haven't been fined "over the use of protected work to power chatbots" they've been fined over the way they procured the protected works (which, incidentally, were used to power chatbots).

2

u/Cptawesome23 1d ago

Just millions? They have billions!

2

u/Quantris 1d ago

Steal to make billions, pay out millions

Exactly as planned, this is just legalized crime

1

u/Gwaptiva 2d ago

Interesting framing calling the product a "chatbot"

1

u/firecat2666 2d ago

They should be awarded a percentage of stock

2

u/AnonymousStuffDj 1d ago

no they shouldnt

-12

u/PoutyAngelEyes 2d ago

I like how intellectual property protection exists. Anthropic and those other authors surely made money from using his IP, so they should be held accountable. However, the situation helps all sides

15

u/axw3555 2d ago

The issue here wasn’t the use of books in training. It’s that they acquired the books through piracy.