AI Companies Hope the Government is Stupid
What distillation is, how it relates to copyright, and why I feel no pity for any AI company complaining about it.
Recently, Anthropic complained yet again about Chinese companies 'brazenly' and 'illicitly' distilling their models. They called it a "distillation attack" to make it sound big and scary.
Now, I regularly tell people that I think Claude is the most capable LLM on the market today. But Anthropic is one of the worst offenders when it comes to peddling fear and doom to promote their products.
In this post, I'm going to break down what distillation is, how it relates to copyright, and why I feel no pity for any AI company complaining about it.
What is Distillation?
I'm going to try to keep this relatively simple instead of going hard-core into the details.
AI models must be trained, and that training is very expensive. If you follow the AI space at all, you know that most AI companies have several different "tiers" of models.
In the case of Anthropic, that's Fable, Opus, Sonnet, and Haiku in order of most powerful to least. Once you've trained a flagship model like Fable, it's common to train smaller models from its output, because it's far cheaper than starting from scratch. This is what distillation is: training smaller models from larger, more advanced, and generally more capable models. It improves the quality of the smaller models and saves money. Win-Win.
What Anthropic is accusing the Chinese firms of doing is spinning up tens of thousands of chat sessions, using the outputs to train their own models on the cheap, and then attempting to undercut Anthropic in the market because they have lower training costs and can afford to sell access more cheaply than Anthropic can.
The Absolute Hypocrisy
Anthropic and many other hyperscalers have been, or are currently being, sued for the massive amount of copyright infringement they have engaged in. For example, Reddit is suing Anthropic for scraping user data from its site in violation of its terms of service. They also bought millions of physical books and scanned them all to feed into their models, arguing that since they paid for the books, it's "fair use" to copy all the information from them and resell it. That didn't get them into legal trouble (though I think it should have), but they were slapped for downloading copyrighted materials from an internet archive without paying.
AI models are data-hungry, and pretty much every hyperscaler company (Microsoft, OpenAI, Google, Meta, Anthropic, etc.) has been sucking down every bit of data they can find, legally or not.
Copyright is a sensitive issue for me. I grew up in the 90s and early 2000s, when the RIAA was suing Grandmothers because their Grandkids downloaded a few songs online. Seeing big tech firms steal everything that isn't nailed down with almost no repercussions makes me feel like I fell into an alternate universe.
In fact, a Reddit co-founder, Aaron Schwartz, literally committed suicide while being prosecuted for stealing millions of scholarly articles.
So seeing any of these firms, who built their entire products on information they didn't own, now whining about Chinese companies doing it to them is peak hypocrisy.
Why This is a Big Deal
Copyright isn't an easy topic for people to understand, but as a creator, it's something I've had to familiarize myself with over the years. In short, if you create something, you automatically own the copyright unless you assign that right to someone else. If you work as an employee or contractor for a company, you'll often have to sign something stating that anything you create belongs to the company that paid you. This is normal and fine. If anyone other than the copyright holder uses the content, they may be liable for infringement and, if found guilty, face financial penalties.
There is a bit of a loophole in copyright, though, called fair use. In copyright law, if someone uses your works in a way that is "transformative", which is a nice vague term, then it is allowed. This is why YouTubers can post "reaction videos," and critics can post clips of copyrighted music and movies in their reviews. There are various legal tests for this, but the main one is that they aren't competing with the original work. Meaning that a movie critic isn't providing an alternative product that competes with the movie; they're just reviewing it.
AI Summaries as Competitive Products
This is the first area where AI companies are running into legal problems. When search engines like Google scrape a website, index it, and display snippets from the site in their search results, they're not transforming the original work. However, they're not competing with the site because users still have to click through to get all the information.
With AI search, we're already seeing data that suggests most users are no longer clicking through to the site. This is bad for the sites because many of them make their money on ads. No user clicks = no ads shown = no money.
In this case, the AI output is directly competing with the original material, and Google isn't paying these sites a dime.
The "Transformative" Argument
When websites started complaining about Google's AI summaries, Google tried to claim that because their LLM doesn't output the original content verbatim, it's considered a transformative work and therefore is fair use. But that opens another can of worms: What happens if the LLM hallucinates?
I'll tell you what happens. AI hallucinations in search summaries mean Google is providing false information about the business. In other words, they're publishing lies. And lying about someone's business and costing them money is something that a lawful society frowns on. And in Germany, a court recently ruled against Google, holding them liable for false statements generated by their AI overviews.
This puts Google between a rock and a hard place legally. If AI summaries are transformative works, then they are legally considered to be statements by Google and subject to libel laws and other liabilities. If they're not, then Google is still in trouble for misquoting a copyrighted source.
Going back to the movie critic. If a critic lies about a film's content, they can rightfully be sued. Fair use doesn't protect them.
Anthropic Can't Have It Both Ways
This is where the fair use argument falls down. AI companies have fiercely argued that their models are transformative works and therefore fair use. But any politician or judge with half a brain is going to ask the obvious question:
Because if model training isn't fair use, then these AI companies owe a lot of copyright holders a lot of money. If the law is to be applied consistently, then either all AI training is free of copyright claims, or they broke the law.
Anthropic is also leaning on its terms of service, which bar the use of Claude's output to train rival models. Maybe that holds up... but there's something rich about a company that vacuumed up the internet, competing with millions of copyrighted works without asking, and now insisting its own fine print is sacred.
It begs the question: if I bury "no AI training" in my site's fine print, is that a contract anyone actually agreed to, or does it only count when a company of Anthropic's size and political connections writes it?
But now that they've complained, they've put themselves in a poor position. Because if they drop the complaint against the Chinese firms, smart people will see that as Anthropic admitting that AI training is not transformative and therefore doesn't fall under fair use.
Anthropic already got into trouble by doomposting about their latest model, which caused the Government to make them turn it off, and this letter-writing campaign could harm them even further.
If distillation is legal, just like training models is legal, then the models these AI companies have spent trillions of dollars training and building out infrastructure for are nearly worthless, because any competitor with the time and compute power to do distillation can release a competing model that is almost as good at a fraction of the cost.
If distillation is illegal, then training AI is not fair use, and they owe us a lot of money.