AI industry horrified to face largest copyright class action ever certified

MicroWave@lemmy.world · edit-2 2 months ago

AI industry horrified to face largest copyright class action ever certified

FauxLiving@lemmy.world · 2 months ago

It looks, to me, like you’re reading the briefing without understanding how the legal system functions. You’re making some incredibly basic mistakes. Copyright violations and theft are two distinct legal concepts, for example. You’re treating the case summary as if it were the legal argument in the brief and you’re misinterpreting some pretty clear legal language written by the judge.

Anthropic admitted that they pirated millions of books like Meta did, in order to create a massive central library for training AI that they permanently retained, and now assert that if they are held responsible for this theft of IP it will destroy the entire AI industry.

No, that is not their argument.

Their legal argument, in the appeal of the class certification, is that the judge did not apply the required analysis in order to certify the three plaintiffs as being part of a class. He instead relied on his intuition, not any discovered facts or evidence. This isn’t allowed when analyzing a case for class certification.

In addition, Anthropic adds, it is well supported in case law (cited in the motion) that copyright claims are a bad fit for class action.

This is because copyright law focuses on individual works and each work has to be examined as to its eligibility for copyright protection, the standing of the plaintiff and if, and how much, of each individual work was the defendant responsible for violating copyright.

This can be done when 3 people claim a copyright violation, because they have a limited set of work which a court can reasonably examine.

A class action would require a court to consider hundreds or thousands of claimants and millions of individual works, each of which can be challenged individually by the defendant.

Courts typically don’t like to take on cases that can require millions of briefings, hearings and rulings. Because of this, courts usually always deny class action certification for copyright violations.

The court, in its order, did not address this or apply any of the required analysis. The class was certified based on vibes, something that doesn’t follow clearly established case law.

Authors concede that training LLMs did not result in any exact copies nor even infringing knockoffs of their works being provided to the public.

This is because training an LLM results in a language model.

A language model is in no way similar to a book and so training one is a transformative use of copyrighted material and protected under fair use.

Authors concede that training LLMs did not result in any exact copies nor even infringing knockoffs of their works being provided to the public. If that were not so, this would be a different case. Authors remain free to bring that case in the future should such facts develop.

In other words, Anthropic can still face liability if it’s trained AI produces knockoff works.

No, the judge didn’t make any claim about the model’s output after training. That isn’t an issue that’s being addressed in this case. You’re misunderstanding how judges address issues in writing.

Here, the judge is addressing a very narrow issue, specifically the exact claim made by the plaintiff (training with copyrighted material = copyright violation).

The subject of the paragraph is concerned with training the LLM. The claim by the plaintiff is that using copyrighted works to train LLMs is a violation of copyright. That’s what the judge is addressing.

The judge dismissed this argument because it was transformative and so protected by fair use.

The judge further noted that the plaintiffs did not show that training the LLM resulted in “any exact copies nor even infringing knockoffs of their works being provided to the public” and if they could show that training the LLM resulted in “any exact copies nor even infringing knockoffs of their works being provided to the public” then they could bring a case in the future. This is the judge hinting that they can amend their filings in this case to clarify their argument, if they had any evidence to support their claim.

The judge is telling the plaintiff that in order to succeed in their claim, which is that training an LLM on their work is a violation of their copyright, they need to show that the thing that they’re claiming has to result in copies of infringing material or knockoffs.

The training resulted in a model. Creating a model is transformative (a model and a book are two completely different things) and the plaintiffs didn’t show that any infringing works were produced by the training and therefore they have no way of succeeding with their argument that training the model violated their rights.

You’re reading a lot of extra into that statement that isn’t there. The plaintiffs never made a claim about the output of a trained model and so that argument wasn’t examined by the judge.