Rendered at 21:00:19 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
himata4113 1 days ago [-]
Does this matter? Distillation is not illegal by every definition of the word.
There are millions of samples available on huggingface and models explicitely trained on output produced by fable. There has been no action taken against them.
Another example is that it appears that the upper limit of what you can do is ultimately dependent on people working on the model, otherwise grok would be a LOT more competitive pre-cursor acquisition.
And lastly, kimi architecture is vastly different than that of fable as it uses mechanisms developed by... kimi themselves. US AI labs are inspired by opensource advancements just as much as open source labs are inspired by traces from models such as fable.
Claiming in any shape or form that fable disillation is one of the primary reasons why kimi k3 is so competitive is slandering the work of other labs that cooperatively push the open-source models forward.
edit: (moved this to bottom)
The only argument they have here is that they use GB300 GPU's which for some reason should not be available to chinese citizens.
jaggederest 1 days ago [-]
Perhaps even more importantly, the current frontier LLM models are self-admittedly the product of enormous quantities of copyright infringement and even less savory inputs, so calling them out for distilling the fruit of that tainted tree reads as highly hypocritical at best.
aunty_helen 1 days ago [-]
No, they settled that yesterday, so all is forgotten. Press releases were queued for today so just in the nick of time.
janderson215 17 hours ago [-]
I think that was specifically on the piracy aspect.
RazorBucksICO 19 hours ago [-]
So Moonshot is next for a lawsuit or do lawyers only care about Anthropic?
anon373839 10 hours ago [-]
Between the two of them, only one firm is trying to pull up the ladder so that no one else can have what it took.
dylan604 1 days ago [-]
This is why I don't give a shit that this is happening. It's actually kind of funny to me.
azinman2 1 days ago [-]
Unless you’re from mainland China, you should.
VulgarExigency 1 days ago [-]
Why? Should we be held hostage to the whims of these companies and all the investors in the throes of AI psychosis, and let them do whatever the fuck they want, because if we don't then the economy will crash?
xbmcuser 1 days ago [-]
No you should not care unless you are a shareholder in one of these Ai ponzi companies. For the rest of the world Chinese companies matching and open sourcing llm's will keep 100 or so tech oligarchs taking over all the world economic output for themselves as all the idiot politician are unwilling to tax wealth.
HDBaseT 20 hours ago [-]
Everyone is technically (directly or indirectly) a shareholder of one of these AI or AI affiliated companies.
cowboy_henk 12 hours ago [-]
We're also shareholders in all other publicly traded stocks (assuming you're referring to pension schemes etc), which means competition in the AI model market is better than a few winners making everyone paying through their nose for access. Even better, open weights models which cannot be disabled on a whim means the entire world economy can benefit.
stuaxo 10 hours ago [-]
The majority being indirect and having no input to how these companies run if we can't get some strong regulation going.
bilbo0s 24 hours ago [-]
Why?
Serious question.
I'm from the US, and I think it's hilarious.
breppp 23 hours ago [-]
One example is that a totalitarian government known for erasing historical record of its crimes will control an arbitrator of truth.
analognoise 20 hours ago [-]
The USA?
sph 14 hours ago [-]
A KGB spy and a CIA agent meet up in a bar for a friendly drink
"I have to admit, I'm always so impressed by Soviet propaganda. You really know how to get people worked up," the CIA agent says.
"Thank you," the KGB says. "We do our best but truly, it's nothing compared to American propaganda. Your people believe everything your state media tells them."
The CIA agent drops his drink in shock and disgust. "Thank you friend, but you must be confused... There's no propaganda in America."
22 hours ago [-]
dylan604 23 hours ago [-]
so what? this has nothing to do with them taking from those that took before them. china censoring information they do not like is nothing new and seems irrelevant to this conversation
breppp 23 hours ago [-]
Fair enough you don't care, some people might want to know about the Uyghur happy camps, mass organ harvesting and such.
In a world where information sources are only going to dwindle, it is not in anyone's interest to empower actors that will use these to manipulate perceptions
VulgarExigency 12 hours ago [-]
The same people who say there is a Uyghur genocide are the ones who deny there is a Palestinian genocide. Guess which one there's video evidence of (a horrific, endless amount of evidence).
dylan604 22 hours ago [-]
The people that care about that won't be using CCP products/services now will they?
4bpp 21 hours ago [-]
And some people might want to know about [insert your favourite beyond-the-pale-in-the-US topic here, we're on a US forum after all]. I think it would be great if those people could turn to Chinese models, while anyone wanting to know about Uyghur camps can ask the US ones.
customguy 20 hours ago [-]
> beyond-the-pale-in-the-US topic
Can you name one? It's an honest question, first of I'm not American, but also the stuff I do come up with (asking an LLM how to blow up a school or whatever) would also be handled similarly in China, so those would be a wash, and I can't think of any that aren't.
4bpp 5 hours ago [-]
I would imagine a lot of things touching upon progressive politics would be affected (there were a few high-profile incidents demonstrating bias like Google's black Wehrmacht soldier pictures, but has anyone rigorously tabulated how the various commercial LLMs respond to questions about the gender binary or heritability of human traits considered good or bad?). Overall, I'm too reluctant to even write out in the abstract sequences of words that I never want to explain to a future job interviewer or HR employee who used GPT-7 to comb the internet for all text that stylistically can be traced to me, but just imagine whatever you believe to be vile and wrongheaded opinions in that general space which it is certainly more than justifiable to prohibit. The things that make you think "banning this is good actually" are exactly the things most likely to be banned (and this is true in China too).
Another thing I would try if I had access to the models and enough proxies to hide behind is asking for advice on software/movie piracy or seeing to what extent the models can be elicited to straight up argue against the validity of intellectual property, though there it seems more probable to me that the US models would be permissive.
antiamerican634 16 hours ago [-]
From an outsiders point of view:
- The Trail of Tears
- The Tuskegee syphilis study
- Use of Agent Orange in the Vietnam War
- Open Air biological warfare testts in civilians eg. in 1950 San Francisco
- The only use of nuclear weapons against civilians?
- Coca Cola and "american culture"
- Neoliberalist economy
- Spreading blame for their sins to other "white" nations
plus one: The text input method to HN comments :(
customguy 7 hours ago [-]
all of these things have books about them published in the US
you are comparing a wooden stick to a fighter yet, try again
quorumsensor 19 hours ago [-]
There aren’t any. The idea that the censorship environment is the same in the US as it is in China is nonsense used to ‘both sides’ away concerns about the CCP.
wasfgwp 13 hours ago [-]
The models to themselves are currently generally not censored though and the guardrails are on the application layer?
Of course that might change in the future but as long as the Chinese companies continue publishing their research and models it only makes it easier for third parties to catch up with them.
queenkjuul 21 hours ago [-]
People in China know about this stuff.
21 hours ago [-]
queenkjuul 21 hours ago [-]
The US already decides what models get released by US companies...
owebmaster 1 days ago [-]
Why so? If I'm from Europe or South America, should I hope Anthropic/Openai win?
m4rtink 12 hours ago [-]
I think it is only rational to want them to fail & to fail hard, to make an example.
To avoid future situations where money is invested on hype only with no regards to what societal disruption it causes.
wasfgwp 13 hours ago [-]
The better question is why would any rational consumer would want anyone to “win”? Disappearance of competition is the worst outcome for them (I guess being squeezed by a monopolistic company from your pwn country is slightly nicer)
RazorBucksICO 18 hours ago [-]
Yes. They’re flawed, but you don’t want a global police state administered by the PRC.
SZJX 9 hours ago [-]
No matter your perception of PRC, I don't see how competition around open-weight models has much to do with a "global police state".
owebmaster 18 hours ago [-]
Who said so? I don't want the current global police state administered by the US government and oligarchies
NuclearPM 19 hours ago [-]
Why?
8note 19 hours ago [-]
fair use, both in the original training and in distillation, or rather, anthropic has no copyright at all over the output tokens
bluegatty 1 days ago [-]
No - distillation is not data inputs.
Raw materials vs. Value add.
They are different things, like ore and metal.
Distillation is a new thing we need to understand, it's probably closer to IP than not.
tikhonj 1 days ago [-]
The "data inputs" were also, very much, somebody's "value added" IP.
We're talking about things like text people wrote, not some kind of raw data floating out in the ether.
bluegatty 1 days ago [-]
Did I say there was no value add in the inputs?
Ore has value, a different kind of value than the output of the refinery.
dijksterhuis 1 days ago [-]
distillation has been around for 12 years. it's not new in terms of ML techniques.
although i doubt there has been a legal case over it yet in the context of the legality of stealing shit but IANAL.
bluegatty 1 days ago [-]
Yes, I get that, but it's only now the issues are coming into the commons in a way that industry / society needs the regulatory clarity.
It's completey insane that we still don't know how Open Source would work, that the laws are vague and we're still technically waiting for the courts to decide on cases.
The government should a) legislate and b) create test cases and run them through the courts so that we can have clarity.
dijksterhuis 24 hours ago [-]
> it's only now the issues are coming into the commons in a way that industry / society needs the regulatory clarity.
you mean like the regulatory clarity surrounding stealing shit to make the LLMs in the first place?
> [There is] extensive litigation on the limits of Fair Use to AI development. Currently, we only have 3 first instance decisions out of the 53 cases being tried. It will likely take a decade before we understand how Fair Use applies to any one step in AI training, let alone all.
> The government should a) legislate and b) create test cases and run them through the courts so that we can have clarity.
if so, it would be nice if they approached the instances of stealing shit chronologically. but that's just my view.
bluegatty 23 hours ago [-]
It's hilarious to see how all these HN intellectuals, who on any other day are rabidly against most DRM / IP issues all of a sudden become private property absolutists.
It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ...
... but Chinese SOTA foundries directly using distillation as fair game.
I don't think there is any coherence to any of these arguments - other than 'we liking big companies'. That's the only common thread.
What is more reasonable:
- There's some grounds for fair use by SOTA models to ingest content, so long as they are not reproducing it ... very roughly speaking.
- SOTA makers are producing novel works, there is value add in that process, again roughly speaking.
- Distillation is a bit of a grey zone, producing random content as arbitrary input is one thing, but producing training sets is another. I think there's a coherent line in there somewhere, I'm not sure where it is.
HarHarVeryFunny 21 hours ago [-]
You're being too charitable to Anthropic, and assuming that the way they are abusing the word "distillation" has some real meaning here. It doesn't.
Anthropic's models simply do not give you their reasoning output - they give a sanitized "summary" instead, for this exact reason, so that the output is not useful to anyone who might want to use it for training.
You can't distill what you are not given - simple as that.
Are Chinese using the output of US models to help create some additional training data for their own in some way? Yes - quite possibly (e.g. LLM as judge), but its got nothing to do with distillation.
bluegatty 21 hours ago [-]
I see your 'fine point' but I don't think it holds - 'distillation' is a perfectly reasonable term to describe the process of creating outputs from one model to that expose key training element, to use in another model.
I think where the definition may be be invalid, is in the creation of 'unrelated data sets for training' models, for unrelated issues.
Creating training sets that mach a models core training, is definitely distillation, it does not have to expose the reasoning traces.
Synthesizing data for some arbitrary thing ... I'm not sure that would be the same thing.
It's hard to draw the line.
But the Chinese models are absolutely distilling - and would not be competitive without this distillation.
At the same time, there's a lot of real innovation and regular building going on at the same time over there.
HarHarVeryFunny 21 hours ago [-]
No - you can't distill if what you are given doesn't have the thing in it that you want to distill out of it.
I don't know why it's so important to you to use the word "distillation", but it's the wrong word to use.
BTW OpenAI on twitter also said that Kimi 3 "cannot be explained away by distillation or anything like that". The timeline of how long it takes to train a model and when Fable was released don't even line up. This is just Anthropic as usual trying to manipulate the US government into helping them shut down competition.
bluegatty 16 hours ago [-]
Distillation is absolutely - and uncontroversially - a valid term for what is happening here.
This isn't really a debate, I'm not making a fine point - just check with all of the various defintions of the term.
Moreover - the 'reasoning traces' are not required for distillation at all.
Finally - it's entirely possible for them to have used Fable for later stage fine tuning.
It's fair to be skeptical of Anthropic (and everyone else) - but this is 'distilling'.
HarHarVeryFunny 9 hours ago [-]
Words have meaning - you cant just redefine them because you want to.
Are reasoning traces required for distillation? Well they are if what you are trying to distill is reasoning, such as coding expertise.
Do you need reasoning traces for "LLM as judge"? No, but it would be highly perverse to call that distillation when there is a more accurate name for it - LLM as judge.
If you want to call use of Anthropic's redacted model outputs in any fashion that violates their terms of service (using them them to help develop anything that competes with Anthropic) as "distillation" then I can't stop you, but it reduces their claims to a joke.
Finally, as noted, OpenAI (who are just as anti-Chinese as Anthropic) said that Kimi 3 can't be explained via distillation (even true distillation!!), or even "anything like it". But random internet guy, you, disagrees. OK.
throw10920 17 hours ago [-]
> Anthropic's models simply do not give you their reasoning output - they give a sanitized "summary" instead, for this exact reason, so that the output is not useful to anyone who might want to use it for training.
This is just straight-up factually false.
The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, making it even more valuable.
There's absolutely nothing about the distillation process that requires that reasoning in the first place, either. That's a definition that you made up.
Chinese models are, factually, distilled from Anthropic models. I've personally repeatedly asked several different Chinese LLMs what their name is, and they answered "Claude".
Don't make stuff up to suit a political agenda. It's extremely dishonest.
HarHarVeryFunny 6 hours ago [-]
> I've personally repeatedly asked several different Chinese LLMs what their name is, and they answered "Claude"
I'm curious what you are doing to get them to override their own name that they were trained on and/or have as part of their system prompt?
I'd assume that the Chinese are scraping the internet for training data the same way western companies do, so for sure there will be a lot of AI generated content in their training data - you don't need to be paranoid and assume they must be getting it all direct from Anthropic.
HarHarVeryFunny 8 hours ago [-]
> The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, making it even more valuable.
Useful for what is the question. Nobody is debating whether the outputs of LLMs are valuable.
Given that Anthropic have redacted their true reasoning, and replaced it with a "summary", specifically designed to be useless for distillation purposes, it would certainly be highly ironic if this summary was in fact "even more valuable" for that purpose as you are claiming!
throw10920 8 hours ago [-]
> Useful for what is the question.
Useful for distillation. Any employee at a frontier AI lab will tell you this. This is known in the industry, and it's an open secret that some US labs (OpenAI) distill on the others. Again - don't just make up stuff for a political agenda.
> specifically designed to be useless for distillation purposes
No, it's designed to give feedback to the user, in a way that minimizes its value for distilling. It's still valuable, and so there's a good chance that they'll remove it entirely as a result.
> it would certainly be highly ironic if this summary was in fact "even more valuable" for that purpose as you are claiming!
I did not claim that. Read my comment again:
> The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, making it even more valuable.
Because apparently I have to spell it out:
The output of a reasoning model is valuable, even if it didn't have the reasoning summary. Anthropic's models have a reasoning summary. The reasoning summary makes the output more valuable than if it didn't have a reasoning summary. It does not make it more valuable than having the full reasoning.
HarHarVeryFunny 6 hours ago [-]
> Useful for distillation
Here's the thing: no-doubt a summary, if it at least reflects some of the logic connecting response to request, is better than nothing, so this can still be useful additional training data, but a model trained on it would be learning to generate these summaries, not the original withheld reasoning, so "distillation" seems an intentionally emotionally-wrought way of describing it (the "summary" is generated by a different smaller model - not the one the rest of the response came from).
It does bring up an interesting point though - RL training in general results in "cargo-cult" reasoning - you train a model to follow the steps (mistakes and all - Karpathy) that got to a result, without understanding why they worked. If this training on summaries works just as well as training on detailed reasoning, then it just highlights how having a few breadcrumbs to follow/regurgitate is all that it takes.
At the end of the day, without having internal logits or original reasoning traces, "distillation" (which suggests one model being derived from another) just seems a very manipulative way of describing this. OK, so it's a terms of service violation - a customer is using Anthropic model outputs to help create something that competes with Anthropic, but that's it. They are not copying Anthropic - they are, one assumes, using other models to generate cheap training data that they would otherwise have to pay people to generate.
This is why people are calling out the hypocrisy - Anthropic are apple-pie American innovators when they appropriate other people's copyright data for training, but Kimi are evil communists when (we assume) they use data generated by Anthropic (not even copyright protected) to help train their own.
wasfgwp 13 hours ago [-]
Well there were observed cases of Claude calling itself Deepseek or Qwen. So pot calling the kettle black?
To be fair I find it hard to take this too seriously, shouldn’t it be trivial to just replace “Claude” with any other string in your “distillation” dataset?
throw10920 8 hours ago [-]
> Well there were observed cases of Claude calling itself Deepseek or Qwen. So pot calling the kettle black?
There are open-source Deepseek and Qwen models - "distilling" doesn't involve breaking terms of service or hitting an API because you can literally run local inference or even just inspect the weights directly, and that's intended because they're open source.
It's categorically different for a nation-state to build massive illicit networks of fraudulent identities to do distillation over tens of thousands of accounts to intentionally bypass providers' terms of service, intention for their models, and business model that very explicitly proprietary and not open source.
If Claude did distill on proprietary PRC LLMs - then fine, shame on them - I condemn that and I expect others to do the same. But there are no open-source Claude models. The only way for PRC models to have those responses is if they distilled Anthropic's models from their APIs.
> To be fair I find it hard to take this too seriously, shouldn’t it be trivial to just replace “Claude” with any other string in your “distillation” dataset?
...and what would happen when it read all of the books and articles about Anthropic and replaced "replaced Claude Opus" with "replaced Qwen Opus"? Did you give any thought to this at all before saying it?
wasfgwp 6 hours ago [-]
> It's categorically
The fraud part and using stolen accounts or credit cards or blatantly violating the terms and conditions (i.e. reselling subscriptions not using outputs in certain ways somebody might not like) is indeed categorically different.
Using uncopyrightable outputs of an AI model obtained legitimately to train your model is not inherently interlinked with any of those things. I don’t really see how the model being proprietary or “open” is particularly relevant when talking about the outputs.
Even using the word “distilling” in this case is deceptive and biased. It implies that the Chinese are somehow stealing Anthropic’s models or their weights and somehow directly transforming them into new models. That’s certainly not what’s happening in any direct sense.
e.g. what if I agreed to send all my Claude code session files to Deepseek or whoever? There would be nothing wrong about that since I and not Anthropic own those files and can do whatever I want with them. Using certain different ways to obtain them of course could be highly illegal.
8note 19 hours ago [-]
idk. i think its fair use when anthropic trains off of copyrighted works, and theres no property rights at all related to the model outputs
there's no creative work between the weights and the tokens being made.
whats the big deal if chinese companies sell an exact replica of the model? its a summary of a variety of works of text and images
wasfgwp 13 hours ago [-]
Do you think that Anthropic should own any outputs you generate with their models?
If not there is not there is no grey zone whatsoever.
dijksterhuis 21 hours ago [-]
> it's hilarious to see how all these HN intellectuals, who on any other day are rabidly against most DRM / IP issues all of a sudden become private property absolutists.
for the record, i've always been rabidly pro-copyright since i worked at a performing royalty organization (prs for music) circa 15 years ago, way before i joined hn.
i don't use llms for that reason.
> It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ...
when i see a spade, i call it a spade. just because the US has utterly stupid copyright provisions that are wide open for abuse, i.e. fair use, doesn't mean abusing those provisions at scale is morally acceptable.
> ... but Chinese SOTA foundries directly using distillation as fair game.
two wrongs don't make a right, but the irony is at least something.
> I don't think there is any coherence to any of these arguments - other than 'we liking big companies'. That's the only common thread.
the corpos can get fucked as far as i'm concerned.
> What is more reasonable: ... There's some grounds for fair use by SOTA models to ingest content, so long as they are not reproducing it ... very roughly speaking.
*only in the US.
wasfgwp 13 hours ago [-]
> two wrongs don't make a right
There is no second “wrong” here.
Model outputs are not copyrightable. I think that was already established?
Or do you think that Anthropic should own all the code generated by Claude? Surely that would be somewhat problematic?
If Anthropic feels that some of their customers are breaking their EULA (nothing to do with copyright infringement though) they are free to stop doing business with them. Maybe even sue them in civil court for breach of contract (again nothing to do with copyright infringement though)
dijksterhuis 9 hours ago [-]
there's a difference between a moral wrong and a legal wrong. in a heavily simplified view, moral wrongs are usually decided in the eyes of the victims -- you did bad thing to me so i'm not going to talk to you anymore. legal wrongs are decided by courts -- you did a bad thing so this court has decided you're not allowed to talk to that person anymore.
anthropic are essentially saying in this tweet they believe a moral wrong has been committed against them -- "unacceptable behaviour" etc.
plenty of people have been vocal about the fact anthropic have committed moral wrongs at scale in building the products in the first place, with the question of legal wrongs still being worked out.
so, two moral wrongs. legally, fuck knows.
wasfgwp 6 hours ago [-]
Sure but it’s hard to read what Anthropic is saying in any other way than that they think that it’s morally wrong to engage in any behavior that harms their (potential) profit margins. The exact phrasing is just a way to justify their stance to other people.
I mean you are right in a way of course, it’s just a matter of degree and perspective, though. If one thing is moderately morally wrong and the other is potentially lightly morally wrong I don’t think it’s fair to equate them.
To me the situation is a bit like Google coming out and saying that its morally wrong for someone to build a competing open operating system on top of Android while stripping all Google services and “stealing” their ad revenue. Just seems silly and hypocritical.
bluegatty 21 hours ago [-]
I'm not attacking you, you don't have to defend yourself.
I'm just nothing that HN rhetoric is contradictory.
But this:
"when i see a spade, i call it a spade." -> this is anti intellectual absolutism.
If it were some true injustice, then fine, but that is clearly not the case.
There is ample room to contemplate that even copyrighted works could be considers fair use as training material.
"the corpos can get fucked as far as i'm concerned."
Ok that's fine - but then don't expect anyone to respect your principles if you don't have any other than 'screw that group!'.
I'm sympathetic to it (!!!) - but if we want to call a 'spade a spade' in a legitimate way, then we can do it in consistent and principled way.
throw10920 17 hours ago [-]
> I'm just nothing that HN rhetoric is contradictory.
Periodic reminder that HN is not a collective or a singular entity and is actually a bunch of different people with different opinions. Often the people with the loudest opinions get upvoted to the top - and often the "side" represented at the top is different from thread to thread.
dijksterhuis 9 hours ago [-]
also, human beings themselves can be contradictory. as an example i will happily pirate films/tv shows, but refuse to do the same with music.
dijksterhuis 9 hours ago [-]
[dead]
m4rtink 12 hours ago [-]
I think people just react to the hypocrisy of corporations stamping on people for "copyright violations" only for (often the same) corporations to blatantly obtain any data they can find, totally disregarding any licenses, scrappers overloading web sites or even privacy.
And the result is force feeding an AI slop generator with a subscription while making personal hardware 3x+ times more expensive.
No wonder people are fed up with this behavior.
queenkjuul 21 hours ago [-]
Like most things, i support rights for people, and not companies. Copyright was created for authors, artists, and inventors. Rules for corporations can and should be different.
bluegatty 15 hours ago [-]
This makes little sense, either from either a moral or pragmatic perspective.
You do realize the 'investors' are the one's who 'own' companies and therefore the IP?
blackqueeriroh 18 hours ago [-]
Lolololololololol
One person can a corporation be.
lovich 21 hours ago [-]
> It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ...
> ... but Chinese SOTA foundries directly using distillation as fair game.
As someone who says it’s fair game, it’s less that I’m being hypocritical and more that I don’t care that one thief had their shit stolen by a second thief. I also wouldn’t care if someone distills the Chinese models. It’s just thieves all around and if they want legal protection or moral outrage from the common man then my view is that they should stop stealing first.
wasfgwp 13 hours ago [-]
Why are you calling the Chinese companies “thieves” though?
LLM outputs are not copyrightable (or rather the user is effectively the only one who can own it). It would be problematic if Anthropic owned all the software generated using Claude..
lovich 6 hours ago [-]
They are thieves the same way Anthropic or OpenAI are thieves. Either it’s fair use to learn from this data or not.
If what Anthropic/OpenAi did for training is theft then the Chinese models also are a form of theft. If they didn’t steal then I don’t think the Chinese firms did either.
wasfgwp 13 hours ago [-]
LLM outputs are not copyrightable or rather the user who generated them owns it.
That entirely settles it and there isn’t much else to say about.
If Anthropic feels that other countries are violating their EULA well they are free to stop doing business with them.
spwa4 11 hours ago [-]
The claim these companies make goes much, much further than that.
They claim LLMs "uncopyright" their inputs. So if I take, say, 50 Mickey Mouse comic books, tell ChatGPT to read them and produce 50 "Buster Beagle" comic books that there is ZERO "copyright contamination" and I own those 50 output comics without Disney having any claims on them whatsoever.
Or if I ask ChatGPT to "make a spreadsheet software like Excel, Sheets, Calc, ..." that, again, there is zero copyright claim possible from these people.
It has not been tested, of course.
sho_hn 1 days ago [-]
Are you suggesting data input is further from IP than distillation?
That would stun me, but it's a little hard to read.
mkehrt 1 days ago [-]
Do you think writing books (and Wikipedia articles, and stack overflow articles, and github repos, and, and, and, and ...) is not a value add?? What terrible claim.
remus 1 days ago [-]
While I agree on a moral level, I think there is a distinction to be made. Training a SOTA model takes a huge amount of resources and expertise so the people doing the training are adding a lot of value along the way. I think this is much less true for distillation (which is kind of the whole point).
ed: to clarify, I totally agree that a huge chunk of the value in LLMs is coming from the source material. My point was just that training an LLM takes more resources and expertise than distilling from an existing LLM so I don't think the equivalence between training and distilling is entirely justified.
Bratmon 1 days ago [-]
I like this comment because its argument only makes sense if you assume that the entire world's output of books and art did not require a huge amount of resources and expertise to make, nor did it add any value.
It's the most CS-major take ever!
mapontosevenths 1 days ago [-]
If turning other peoples copyrighted work into a model is transformative enough to be protected then so is distilling that model into a different, better, model.
foo12bar 1 days ago [-]
The models were built using copyrighted works, so why can't models be built using other models?
usef- 23 hours ago [-]
They do seem to be paying for it (as per the 1.5Bil lawsuit yesterday and them now purchasing books and licensing from media companies).
Whether we think they're paying enough is another question, but "I'm paying for content so can protect it" doesn't seem inconsistent.
We may decide that giving models away for free means they don't have to license content (judging by HN comments), but currently that doesn't seem to be the case as Meta is facing lawsuits for its open models.
The judge found their use is fair use. They are paying not for their use of the content, they are paying for using illegal copies of the content.
The same principle can be applied to distillation - it is a fair use. You just shouldn't use illegal ways to access the models being distilled.
To the commenter below: if it is illegal - has the police/FBI report been made? Otherwise it is just a civil court matter.
usef- 23 hours ago [-]
Fair, but isn't "illegal" access what they're talking about in OP?
It does seem to be becoming the norm for AI companies to licence premium content in America, judging by the deals they're making. It doesn't seem to be done by the international distillers. It's a cost that American open models will seem to have to pay but not international.
Bratmon 22 hours ago [-]
> It does seem to be becoming the norm for AI companies to licence premium content in America, judging by the deals they're making.
This is a very surprising claim to me (and I imagine many small website owners who keep getting scraped by Anthropic and OpenAI).
Do you have a source?
usef- 22 hours ago [-]
There's been many news stories of it over the past year(s) as they signed each one. Here's the first result I could see with a rundown of many of them (am on mobile).
The news corp one had a leaked price ($250mill), so they don't seem to be insignificant. These would have to be included in API prices I presume.
trhway 22 hours ago [-]
>It does seem to be becoming the norm for AI companies to licence premium content in America, judging by the deals they're making. It doesn't seem to be done by the international distillers.
International distillers doesn't use that premium content, so they don't pay for it. They do pay for their access to the models they are distilling. Thus providing the revenue stream to those models. Thus those models make profit off the content they used for training. The content they mostly have't paid for.
>It's a cost that American open models will seem to have to pay but not international.
It goes both ways - American companies and their business are protected by American laws and have access to the market protected by those laws, etc.
usef- 22 hours ago [-]
> International distillers doesn't use that premium content, so they don't pay for it.
This doesn't seem to be true. They are training on their own scraped data overwhelmingly (we can extract copyright data from, eg, deepseek). They couldn't get nearly enough tokens through the American APIs to train a model on alone.
> American companies and their business are protected by American laws and have access to the market protected by those laws
Absolutely. Currently international providers are selling inference on the American market though, I don't know how that will sit legally the way things are currently going.
breppp 24 hours ago [-]
Because model output is probably far closer to software or a licensed work which possibly has greater protections than it is to copyright. There is far less possibility of fair use, it might be protected by patents, license or reverse engineering laws.
In any case the laws are being written now, but I doubt these will have worse protection than software does, which has far better protections than copyright
giaour 23 hours ago [-]
> I doubt these will have worse protection than software does, which has far better protections than copyright
Software is protected by copyright. Some software may also be protected by patents, but last time I checked, AI generated output of any kind was not patentable.
trhway 23 hours ago [-]
Distillation isn't a copy. Distillation is more akin to "clean room" implementation.
Also note that the OpenAI/Anthropic argument is that the model training is sufficiently transformative to satisfy the fair use of the original content for training.
By that same argument, when distilling the distillers aren't using the original content the OpenAI/Anthropic models were trained on - the distillers are interacting only with the "sufficiently transformed" content of the OpenAI/Anthropic models and are normally paying for that.
There is also that old phonebook rule that facts can't be copyrighted. So, if i asked the model about bunch of phone numbers, i can publish the resulting list, can train my model on it, etc. Such approach doesn't allow to reproduce copyrighted works of course - and as we know the AI output isn't copyrightable, so it looks like basically any output i get i can use whatever way i like.
breppp 23 hours ago [-]
Software is protected by the DMCA, patents, licenses, EULAs, all of those aren't there for books. I doubt new laws won't be written for model outputs.
Also, if model output distillation is shown as some form of reverse engineering I assume the DMCA can apply
bigiain 22 hours ago [-]
The C in DMCA stand for Copyright. All (I think?) software licenses are underpinned and made legally enforceable by copyrights. EULAs are underpinned by licenses which are founded on copyright. Patents are the only one of those protections that are not based on copyright, and there are lots of very good arguments against at least most software patents (all software patents of the form "Do {well known and obvious thing} with a computer" should, in my opinion, be immediately revoked and potentially have every company who's enforced payments from such patents investigated for fraud).
giaour 23 hours ago [-]
You may recall that the DMCA was originally written to protect music and movies. It does in fact apply to creative works. If you have ever purchased an MP3, eBook, or streaming movie, you will also be aware that you purchased a license to the underlying IP. This is also true of physical media, but the license agreement you have to accept when obtaining a digital work makes this explicit.
I agree that you can't patent a book, but I would point out that you can patent an idea, which may only appear in a book or journal article.
vel0city 18 hours ago [-]
You do patent ideas, but the actual words written in a book describing that idea would only be protected by copyright at best. FWIW, the exact words describing the idea being patented are technically public domain; that's the whole point. You're free to go look up that patent, print it out, make whatever copies of it you want. Take any of the drawings in patents, put them on t-shirts, and sell them. No problem. Implementing the ideas those words represent is a different story.
For example, a patent describing a chemical process. The actual idea of how to do it is public domain, go look up the patent. Print it out. Do whatever with those words. Its fine. Building a plant to go do that chemical process to make that same output chemical in that same way, that's IP infringement. Its not the words, its the idea.
wasfgwp 13 hours ago [-]
How is “model output distillation” different to using outputs (which are legally copyrightable) for any other purpose?
queenkjuul 20 hours ago [-]
Afaik (and ianal) there's nothing stopping anyone from attaching a EULA to a physical book
preg_match 23 hours ago [-]
Why would this be the case. Why would software output from a model magically have greater protection than the software the model trained on.
vel0city 23 hours ago [-]
Let's assume model output can be claimed by copyright or some form IP. You can't really patent it, as the output isn't a novel idea or process, much like you don't patent a book or a movie. But for arguments sake, let's agree it is some kind of IP.
Who are you saying owns that IP? The people who trained the model? The people who ran the model? The people who wrote the prompt? The person who paid for all of that to happen?
If the model output is owned by the person prompting it and paying for the tokens, what's the problem here?
If the model output is owned by the trainer of the model, that's a big nasty can of worms.
wasfgwp 13 hours ago [-]
LLM outputs are not copyrightable. At least that’s the current established legal precedent in the US. The only question is whether the user owns the copyright without significantly transforming the output but that’s not really relevant in those specific situation.
I mean otherwise it’s a very slippery slope, effectively it would give Anthropic the ownership of any code generated by its models..
arbitrary_name 23 hours ago [-]
there is a major god complex here.
MBAs and non technical managers = inept Catbert-type charlatans.
Software engineers, devs, etc = geniuses capable of mastering any domain, innate ability to be right on any topic.
joshuamorton 1 days ago [-]
I don't think that's what it's saying at all. It's saying that there's a level of creativity in model creation that isn't present in distillation.
skybrian 1 days ago [-]
Maybe, but it's not like their AI is likely to repeat it back verbatim so it's unlikely to be a copyright violation. It seems like at most, they would be breaking Anthropic's terms of service?
Or maybe they're going through an intermediary "transfer station" that's breaking terms of service:
Yes, it's just a ToS violation at present. Those are legally binding though, despite the common adage. What that really translates to here though, anyone's guess.
Anthropic's own copyright infringement could apparently be forgiven for 1.5B USD after all, so maybe there's a price that breaking the distillation clause for is acceptable too. Or some other arrangement.
Bratmon 23 hours ago [-]
But surely at least one of the websites Anthropic scraped to make Claude had a ToS forbidding automatic access?
Why is Anthropic's ToS any more binding than that of a rabidly-anti-ai literature blog with 50 readers?
skybrian 20 hours ago [-]
One reason is that they might not have scraped it themselves, so if there was a ToS, it was someone else who broke it. For example, see:
There is an even higher level of creativity in creating books, songs and all sorts of art used in model training though. That's your apparent blindspot.
There is no world in which me vacuuming the entirety of human knowledge to make a genai model is ok but hoovering my model answers is not. The hypocrisy is stunning and risible.
Now if you go and make a model based on purely synthetic data and not a single work made by humans, you would have a valid point.
remus 15 hours ago [-]
> There is an even higher level of creativity in creating books, songs and all sorts of art used in model training though.
No argument here, I completely agree.
> There is no world in which me vacuuming the entirety of human knowledge to make a genai model is ok but hoovering my model answers is not.
I disagree with this though. Clearly LLMs owe a huge debt to everything that has come before, but surely you'd agree that the models that are produced are something substantial and new and novel which didn't exist before and have lots of value in their own right. Let's be a bit reductive and pretend Moonshot had just outright stolen the weights from Fable somehow, clearly that wouldn't be contributing anything really new or novel. Now of course they've distilled rather than stolen, but the point is similar: how much value have they added along the way?
gozucito 8 hours ago [-]
Since this is HN Think of it like one of the GPL license for software.
It's ok for me to use your source code for free as long as I then let others also use my source code for free.
joshuamorton 23 hours ago [-]
So, I'm not the person you were responding to. I'd like you to take a moment and suggest where anyone in the thread you're replying to, either me or Remus, has said anything that suggests disagreement with the statement
> There is an even higher level of creativity in creating books, songs and all sorts of art used in model training though.
He claimed there was more creativity in model training than in model distillation. That makes no claim about the relationship between the creativity in model creation and art. Why are you continuing to attack a claim that was never made, after a sub thread very explicitly clarifying that that claim was not made?
gozucito 22 hours ago [-]
This is the post Nemus was replying to:
>Perhaps even more importantly, the current frontier LLM models are self-admittedly the product of enormous quantities of copyright infringement and even less savory inputs, so calling them out for distilling the fruit of that tainted tree reads as highly hypocritical at best.
Context is important. And in this context, their argument only mentions creativity when it belongs to an AI lab. That omission is the blind spot I pointed out. Bottom line is whether or not Anthropic are being hypocritical and yes, they most definitely are, regardless of any attempted sophistry.
There is a reason courts want you to tell "The whole truth" and not just "the truth".
trhway 23 hours ago [-]
>a level of creativity in model creation that isn't present in distillation.
the same argument - a level of creativity in the world knowledge creation that ins't present in the model training on that knowledge.
Or in other words - model creation and training is just a distilling of the world knowledge.
joshuamorton 23 hours ago [-]
I don't disagree. I'm not sure why that's a relevant reply though.
If you think that the addition of a less creative process (model creation) to a more creative corpus ("art") is problematic, then it follows that you should think the addition of a less creative process (distillation) to a more creative corpus (a model) is also problematic.
trhway 23 hours ago [-]
I think both are natural and fine. Otherwise we'd have to outlaw analytical thinking.
perching_aix 1 days ago [-]
No? They outright say the opposite!
Like look, I'm not a native speaker, sure. But I think when someone says "value add", that means there was value there (which you claim they're rhetorically erasing), and then that was added to. Under no interpretation of this phrase do I get an erasure of prior value.
So certainly, as long as words mean anything, no, they absolutely did not say or suggest what you claim they did, and what you extract a thus unreasonable amount of obnoxious schadenfreude from, while throwing in a cheap insult for funsies at the end.
The LLM output, is not the same as the input - there is value add.
Of course works used as raw inputs to LLMs required work and are reasonably subject to IP concerns - but they are different.
It's possible that the LLM makers 'owe' the content creators that created the content they used to make their products - it's an interesting but separate question.
We could very well end up where content IP is protected, LLM output is not and visa versa with reasonable legal founding, doubtful but plausible.
blks 1 days ago [-]
Lossly storing IP in LLM itself, and using IP for training (so it’s lossly stored in LLM), without licensing these works or otherwise following license agreements (eg GPL) is infringement. Using then this product for commercial activity is a smoking gun.
bluegatty 1 days ago [-]
"Lossly storing IP in LLM itself, a" - that part I'm inclined to agree with.
But it's debatable if that's the case.
Google stores copyrighted content and produces in in their product.
Also - it's fair game to use snippets of things here and there, if the derived work is novel, which I think it is for LLMs, mostly.
I do agree though, that we ought to draw the line somehow.
jaggederest 1 days ago [-]
> but they are different.
How, and why?
> We could very well end up where content IP is protected, LLM output is not and visa versa with reasonable legal founding, doubtful but plausible.
That is the current state of legal rulings - LLM output is public domain, not copyrightable.
semiquaver 24 hours ago [-]
This misstates the small number of legal opinions and orders on this topic, none of which form binding precedent outside the districts where the cases happened. So even if a court had found that “LLM output is public domain” (none did) that wouldn’t make it “the law” until it went up the appellate system and was upheld.
Our current laws simply weren’t built for this and I expect the legal status of LLM output is not going to be resolved until Congress actually legislates on this topic.
bluegatty 1 days ago [-]
"> but they are different.
How, and why?"
How are they even remotely the same?
They're not even used the same way.
One is raw data input, the other is training content - designed to train LLMs.
One is a set of IP derived for other purposes entirely, and has esablished IP law - how you can use someone else's creative work or not ... for LLM outputs, less clear.
skippyfish 1 days ago [-]
> raining a SOTA model takes a huge amount of resources and expertise
Writing books, building Wikipedia, and answering questions on online forums takes a lot of resources and expertise that scraping didn't. So at the very least, we're already one rung down the "maybe you should've asked" ladder.
jaggederest 1 days ago [-]
I suspect that, in aggregate, all of the informational output of humanity prior to 2020 has taken more resources to produce than the last few years of LLM research.
ryandvm 1 days ago [-]
I don't know man. This reads like "yeah we stole your grain, but making bread is hard."
Muromec 1 days ago [-]
It sure is, but it doesn't matter. Whatever position that generates more economic activity is declared legal using some nonsense retconned logic "because we said so".
il 1 days ago [-]
Probably not as much effort as writing books and creating art the models were trained on.
GTP 23 hours ago [-]
Still, AFAIK Kimi's architecture (just like that of other LLMs from Chinese labs) is different from those of OpenAI and Anthropic's model in a nontrivial way. So the expertise is still there, and I guess resource use too (although Chinese labs tend to optimize this, thanks to the restrictions they have on GPU use).
EDIT: just wanted to add that resource optimization is usually where the contribution of Chinese labs is, so you shouldn't reaad the above parenthesis as a negative comment.
InsideOutSanta 1 days ago [-]
As an author, that's a genuinely disheartening thing to read.
It took me a year to write a book. It took OpenAI and Anthropic a fraction of a second to ingest it. Do you understand now why I give zero shits if it takes Anthropic a billion to train a model, and Moonshot 10k in API cost to distill it?
1 days ago [-]
liuliu 1 days ago [-]
> training an LLM takes more resources and expertise than distilling from an existing LLM
This is not automatically true. Training and distillation use the same underlying infra and method and there is no intrinsic differences in between.
blks 1 days ago [-]
They add value on top of other people’s work, often against licensing, and then commercialize this product, ie profiting from making a product out of other people’s IP.
altmanaltman 1 days ago [-]
Why is it less true for distillation? Everyone technically has access to Fable but Moonshot came up with the model. How can you objectively claim one is adding value while the other is not?
If that is the whole point you need to clarify why this is the case on an objective level.
I would say building a comparable model using any means necessary (just like what Anthropic and OAI did) at a lower cost is actually more valuable to soceity and Monshoot is arguably generating more value with less.
1 days ago [-]
AlienRobot 1 days ago [-]
The value of LLM's come from replacing what generated its training data.
If the distilled model is cheaper, then it's just LLM's getting LLM'ed.
mort96 1 days ago [-]
Yeah, there's a difference. One party spends a bunch of resources doing something illegal and extremely immoral. The other party spends little money doing something legal and morally neutral.
darod 1 days ago [-]
You can argue that reverse engineering anything is as hard if not harder than engineering something. I can’t imagine distillation is any different.
some_random 1 days ago [-]
Distillation is objectively easier than training a model from scratch, that's why all these Chinese labs are doing it.
recursive 24 hours ago [-]
Training a model is objectively easier than generating the sum total of human creative output prior to 2020. That's why the big labs are doing it. What's the difference here?
breppp 23 hours ago [-]
Reverse engineering has stronger protections than merely copyright infringement
Teever 1 days ago [-]
I'm sure it takes a lot of time and resources to plan and pull off an epic heist but it is unusual to see people like Thomas Crown being accused of creating value, as they're usually accused of committing theft.
Barrin92 23 hours ago [-]
>Training a SOTA model takes a huge amount of resources and expertise so the people doing the training are adding a lot of value along the way.
producing the entire body of human knowledge that Silicon Valley companies absorbed like the Borg did not just take more resources but also a fair amount of blood and sweat, certainly more than the LLM so on that front that comparison also seems entirely justified.
petilon 1 days ago [-]
I disagree that LLM models are the product of enormous quantities of copyright infringement.
The recent announcement that AI-assisted research produced a counterexample to the Jacobian conjecture--a long-standing open problem in algebraic geometry--shows the original value AI can create. The result was not copied from a textbook; it emerged from AI learning from existing material, much as a human does, and then applying that knowledge in a new way. If that's a violation of copyright, then a human doing the exact same thing would be a copyright violation too. But it isn't.
InsideOutSanta 1 days ago [-]
If you re-read your comment, you will find that your second paragraph is not evidence for the claim you make in your first paragraph. In fact, your first paragraph is just false.
petilon 1 days ago [-]
Let me explain it this way: If it is legal for a human to learn from a book, then disseminate the knowledge, then it is legal for a machine to do so. You may think this is not right because a machine does it at a much larger scale, but if so laws need to be updated. As it stands now there is no law that says if a human does X it is not a copyright violation but if a machine does the same X it is copyright violation.
worik 24 hours ago [-]
> If it is legal for a human to learn from a book...
True, if the human's access to the book was legal
A great deal of training was on the open web, no one should complain.
But at least Meta and Anthropic were caught red handed taking copyrighted works, illegally, for training
I think international IP laws are too strick and onerous, but they were broken to train these models
butlike 1 days ago [-]
Yup if a machine kills a human it's not the machine's fault; it's the human's. Humans doing the exact same thing as machines aren't 1:1.
petilon 1 days ago [-]
If it is legal for a human to do something then it is legal for a machine to do it too. Are there any counter examples to that?
yencabulator 23 hours ago [-]
Get elected president?
fooofw 23 hours ago [-]
Lethal self defense?
GTP 23 hours ago [-]
The GDPR gives you the right not to be subject to automated decisions, so there are cases where a human can make a decision and a machine cannot.
ilovecake1984 1 days ago [-]
The didn’t pay for the books.
It’s massive copyright infringement.
The human buys the books.
petilon 1 days ago [-]
Did they borrow the book? If I learn from a borrowed book is that copyright infringement?
foxglacier 24 hours ago [-]
It's worth keeping in mind the purpose of copyright. It's a pragmatic tool to encourage investment in creative work for the benefit of everybody/consumers. We may be entering a time where there's less need to incentivize people to write books. At least not non-fiction books which are simply a collection of existing knowledge presented in an a way that's suitable for human readers. A lot of the value those authors provided can now be done by AI. Yes, the AI trained on their work, but now that it's here, we don't need new non-fiction authors quite as much as we used to.
I wouldn't want to live in a world where technology or general people's wellbeing was held back by obsolete laws that ended up lingering on just to protect undeserving special people at the expense of the rest of society. Remember guilds for tradesmen? They were also a monopoly given by the government to special people. They had their purpose but nowadays we have different ways to keep tradesmen working effectively like license requirements and insurance.
Just to be clear, I think we do still need copyright, but that we might be in a transition period where it has to be redesigned to adapt to AI.
ilovecake1984 22 hours ago [-]
I don’t think it’s been common to write non fiction for money for decades. What they are doing is killing off the real motivation to do it, which is recognition and attribution.
We will all be sorry when professionally written and edited works disappear. An author has a reputation and the incentive to protect that reputation keeps standards high.
ryandvm 1 days ago [-]
Boy I tell you, I am having an awful hard time summoning pity for the organizations that have themselves distilled all of humanity's knowledge into mysterious labor-market-masticating black boxes.
> protect our first-party products from abuse like bots, scraping
Won't you think of the trillion dollar corporations?!
atleastoptimal 1 days ago [-]
It matters because everyone imagines the inevitable "closing of the gap" between closed and open source, but the rate at which open source catches up with closed source seems to depend on being able to train on and distill the outputs of open source models. As long as performance of open source models is at least partially dependent on frontier-model outputs, then that gap will remain in place by definition.
>Claiming in any shape or form that fable disillation is one of the primary reasons why kimi k3 is so competitive is slandering the work of other labs that cooperatively push the open-source models forward.
If the distillation is irrelevant to why it is competitive, why do they do it then? Obviously is helps improve their benchmarks/performance to some degree, otherwise they wouldn't need to do it.
himata4113 1 days ago [-]
Never claimed that it is irrelevant. And kimi k3 is on the same level and sometimes outperforms fable 5 - that cannot be explained by distillation. The reason why gap is not closed is simply the fact that fable was trained months ago so in theory the frontier labs are still 1 (small) step ahead.
Although I will reiterate the fact that distillation is not the primary reason why these models are performing so competitively.
atleastoptimal 1 days ago [-]
Closed-source models have to deal with the current frontier being heavily regulated. Fable, at its old level, was "too good" to be released and they had to add an additional safety layer to sanitize the outputs. Lowering the quality of the models so they are safer and more steerable has been something all the closed-source models have been doing for a while, a requirement that many open source models don't need to deal with.
If Kimi k3 really were above Fable 5 then there invariably the USG would have to consider their restrictions on model capabilities excessive, or one would have to admin closed source models are held to more restrictive safety standards than open source models.
>Although I will reiterate the fact that distillation is not the primary reason why these models are performing so competitively.
How would you know this? How could you ascertain exactly how much performance is attributable to their unique engineering/research? If they really were so competitive they could surely make a model that isn't dependent on distilling Fable or other frontier models.
himata4113 1 days ago [-]
I recommend reading some of their research it's honestly astonishing how intelligent some of their solutions are.
Kimi specifically relies heavily on reasoning traces which is largely due to their training strategy and will perform poorly when thrown into a conversation from another model. Another fun advancement is that they simply ctrl+c ctrl+v'd attention which means that the model can steer where to look in the context window without ever producing an output token increasing token efficiency and attention accuracy as a side effect you end up with weaker prompt adherence.
None of these 'issues' manifest in US models which proves that kimi has diverged and is achieving these capabilities seperately from the architecture that US labs rely on.
I would agree with you during the Deepseek R1 era, but US labs were heavily inspired by open research at that point as well so I wouldn't give them too much credit.
spwa4 1 days ago [-]
You still left out that it doesn't matter anymore, just like Anthropic/Facebook/OpenAI only really needed to read massive amounts of copyrighted data only once (and of course, they all did this illegally, which makes their current complaints more than a little ...). Once they have a large model trained on the data, they can just retrieve reasoning traces and copyrighted data from the previous model. In fact that is a training technique long used because it has better results that directly training on the original data.
In other words: even if the US (somehow) denies them access to the current OpenAI/Anthropic models, they'll be able to improve based on what they already have.
kevinqi 1 days ago [-]
I agree distillation isn't illegal; I also think Moonshot/Kimi is very impressive. But the more interesting question is whether labs like Moonshot can be a real competitor to OpenAI/Anthropic. If you can only play catchup (however quickly you do that), then you're never going to be at the frontier - I think that's why distillation matters.
mring33621 1 days ago [-]
People that think the Chinese are only able to copy western tech are in for a wakeup call.
Actually, that has already happened in many domains, it's just that most western people (USA especially) won't admit it.
kevinqi 1 days ago [-]
my assertion isn't that china isn't able to surpass western AI. I think it may well happen. I've been to china many times and am well aware of how ahead they are in many technological/societal areas.
at the same time, I don't buy the idea that distillation is unimportant in assessing what Chinese labs are capable of. If it wasn't, why did Kimi's release timing coincide so well with Fable's launch?
and if Anthropic hadn't released Fable, would we have Kimi today? If the answer is no, then I think that's still a very important point to consider.
ericmay 1 days ago [-]
Spot-on and you're asking the right questions. And the other problem in these comments is that folks seem to think if China pulls ahead we can't just distill their models, provided that distillation is a key part of "this". If it's such a great strategy we'll just use it too if we want to. Boom roasted.
For some reason folks seem to think that China can take action and then other countries can't also take action or respond to that action and it comes up again and again. China has hypersonic missiles! Pack it up boys time to go home. Nothing we can do. Dang shucks. China distilled American AI models, welp time to just close it all down and let's just write off those trillions of dollars and all the literal geniuses financing and building these things. Oh well China can just copy American models while we spend all the money! Ok we just stop developing models and we'll just copy their models. China will flood the market with their cheap products! Nope can't do anything like, oh, idk, not buy any of those products or just raise the prices on them in local markets. It's never-ending. I don't understand the lack of capacity to reason about other actors that takes commonly takes place. And that's just China, never mind other general issues.
Daishiman 1 days ago [-]
> and if Anthropic hadn't released Fable, would we have Kimi today? If the answer is no, then I think that's still a very important point to consider.
That works both ways, competition and performance spur new developments. You don't think the American labs are looking at Chinese research on how to reduce compute per token?
nylonstrung 1 days ago [-]
So many of the breakthroughs and architecture that make LLMs powerful in general today came from China, especially ones related to sparsity and MoE that have made inference and training substantially cheaper.
Let's not forget how much people talked about "prompt engineering" before Deepseek mainstreamed the idea of thinking mode which is now universal
overgard 1 days ago [-]
I think it depends on where you think we are on the S curve of intelligence growth. (Yes, I think it's an S curve, not an unbounded exponential). If you think we're near the peak than playing catch up (especially if you can play catch up quickly) is very rational.
I know this isn't exactly a scientific test, but I had a local Qwen 3.6 27B model implement a fairly sizable feature today. There were a couple of bugs, mostly around me not giving sufficient specifications, but they were ironed out quickly when I pointed it out. I was able to ask the model to create instructions so next time it doesn't fall into the same pitfalls, and it did a great job. 27B local model! (And it was super fast too).
I ran Fable 5 as a code review and it didn't really have any significant corrections.
I guess my point here is that, for most work the frontier models are probably overkill anyway, and improving on overkill in a way that raises prices significantly is probably not a winning strategy.
The only place I can think of where the super high powered models are "required" is if you want to do a ridiculous token burn like GasTown where you just have it run un-monitored on very long tasks. To me though, that's an experiment, not a real workflow. And the way these labs are like "oh we made this (broken) thing in a week using just agents!" always also follows with "and it cost $100,000+ in tokens!". Like, ok, I get it if you're doing research but that's the salary of an entire person.. that can actually learn and improve.
overfeed 23 hours ago [-]
> But the more interesting question is whether labs like Moonshot can be a real competitor to OpenAI/Anthropic.
The answer depends on whether you think the AI researchers at Chinese labs are (or can be) as smart, motivated, and as good at math as those working at US labs - a not-insignificant proportion of whom are Chinese nationals.
himata4113 1 days ago [-]
My entire point was that this was not achieved purely from distillation and claiming that is slander against open research.
mNovak 1 days ago [-]
> The only argument they have here is that they use GB300 GPU's which for some reason should not be available to chinese citizens
Note that Chinese companies are free to rent from GB300 clouds internationally. There are large datacenter hubs in Singapore and Malaysia serving chinese and other customers.
Though there is also reported [1] significant smuggling of Nvidia chips into China as well.
It is incredibly important to whether the US can maintain its AI lead. If foreign competition is closing the gap only by distillation, then the frontier labs can focus on preventing distillation and maintain their lead that way.
US dominance is also important for approaches to safety, especially political approaches. If the frontier models are all US-based, safety might be tackled via internal US policy. If other countries can independently train competitive models, international cooperation is required.
Edit: It is also important for the business model. Companies won't be able to justify tremendous training costs if competitors can replicate their product much more cheaply via distillation.
titanomachy 1 days ago [-]
Do Americans even believe that US policy is likely to steer development in a way that’s safe and beneficial for humanity? The rest of the world certainly doesn’t. The US currently seems to primarily use their superpower status to be the world’s number one shit disturber and geopolitical antagonist.
I don’t think China’s necessarily any better, but I’d rather have the most powerful models be open rather than under the exclusive control of the US executive.
Gajurgensen 6 hours ago [-]
I didn't mean to imply that the US is more likely than elsewhere to responsibly steer AI via policy. But I do think it is easier if it can be done internally as opposed to via international dealmaking.
CuriouslyC 1 days ago [-]
China uses its power to make favorable deals and get people hooked on what it's slinging so it has captive customers. The US uses its power to bully and break rules that apply to everyone else for its own benefit. Kind of a big difference.
thesmtsolver2 19 hours ago [-]
Not really. Go ask someone in Tibet/HongKong/Taiwan/Japan/India if China isn't breaking rules and bullying them. If China had US's powers/economy, it would be a much bigger bully.
> Do Americans even believe that US policy is likely to steer development in a way that’s safe and beneficial for humanity?
No, at least not outside of this forum.
We all mostly think these models and US policy are going to drive the exact opposite of that. Wealth will continue to get extracted and funneled to the top, and the rest of us are going to be left with the scraps and left to die while what little social safety nets we had continue to get eroded away alongside losing our jobs.
Gallows4574 1 days ago [-]
>Do Americans even believe that US policy is likely to steer development in a way that’s safe and beneficial for humanity?
No, no we do not.
realusername 1 days ago [-]
> If foreign competition is closing the gap only by distillation, then the frontier labs can focus on preventing distillation and maintain their lead that way.
We already know it's false because you would have hundreds of competitors if it was that easy.
The reason why these Chinese labs are releasing good models is simpler, they have access to a tremendous pool of talented people.
slibhb 1 days ago [-]
Of course it matters. Regardless of whether distillation is legal, there is a difference between training a model with and without distillation. For one thing, the distilled model wouldn't exist without the model it distilled.
Also, companies that use distillation may be competitive but seem unlikely to surpass the companies that are training these models from scratch.
bilbo0s 23 hours ago [-]
>but seem unlikely to surpass the companies that are training these models from scratch
Then why is it a problem?
Another serious question.
Trying to get my head around what the root of the objection is here. There must be some fear, but if that fear is not a fear of being surpassed in the market, then what is the fear?
slibhb 22 hours ago [-]
I didn't say it was a problem, I said it "mattered"
If I wanted to argue that it's a problem, I'd just say that companies investing billions in training frontier models should reap the rewards. And distillation is essentially theft.
overfeed 23 hours ago [-]
A fear of competition causing a failure to recoup the trillions invested in AI via sky-high margins, and starting a (short) chain-reaction that causes the bubble to pop (or fizzle). A lot of people have a lot riding on the AI bubble not popping.
JKCalhoun 1 days ago [-]
Legal, illegal…
The word I would use is inevitable. It reminds me of the (PC) clones wars…
thewebguyd 23 hours ago [-]
> Claiming in any shape or form that fable disillation is one of the primary reasons why kimi k3 is so competitive is slandering the work of other labs that cooperatively push the open-source models forward.
They claim it because Anthropic are planning to push for protectionism. They just doubled their political spending to $40 million for the midterms to "push for AI regulation" Gee, I wonder what it is they are lobbying for. Certainly won't be OFAC sanctions right? ICTS import controls?
US GOV, under lobbying pressure from Anthropic and OpenAI are going to go full protectionism and restrict Chinese models, I'd almost be willing to bet money on it. They can't really enforce for individuals, but they can definitely tell US based hpyerscalers they can't host them, make it illegal to host the weights, and government procurement restrictions.
torginus 23 hours ago [-]
> The only argument they have here is that they use GB300 GPU's
I don't follow events closely, but the US has constantly flipflopped on what sort of GPUs the Chinese are allowed to have, not in small part because much of the AI boom's valuation is based on demand for US-made hardware, for which the Chinese have inexhaustible and well-financed demand.
So even this feels a bit hypocritical to me, but my understanding is that Chinese native AI hardware is getting good enough that labs dont feel a huge disadvantage by being forced to buy at home, even if they'd have preferred to buy US chips.
Which is a situation that was manufactured by the constant thread of having their access to advanced GPUs revoked.
pgt 1 days ago [-]
It matters because it means that lab could not train that model without distilling another frontier model, and their progress would slow once they get properly cut-off. If I funded that lab, I would want to know that.
GuB-42 23 hours ago [-]
> Distillation is not illegal by every definition of the word.
I am waiting for a precedent on this one. In general, training on copyrighted material is legal, there is a lot of precedent there. But every now and then there is a case where the owner of the training material wins.
I don't remember the details but I believe one of these instances was when one company trained its AI on the knowledge base of another company and turned it into a competing product. Fair use was denied because of that direct competition. Distilling a LLM to make a competing LLM looks kind of like this, or maybe not, I don't know.
It would make sense for distillation to be legal in every way, LLMs are built on a broad interpretation of fair use, but sometimes, law is weird.
unknownfuture 21 hours ago [-]
> I am waiting for a precedent on this one. In general, training on copyrighted material is legal, there is a lot of precedent there. But every now and then there is a case where the owner of the training material wins.
You're making a fundamental assumption: that model outputs are subject to copyright. In the US that's only the case if a human is part of the creative process:
> It concludes that the outputs of generative AI can be protected by copyright only where a human author has determined sufficient expressive elements. This can include situations where a human-authored work is perceptible in an AI output, or a human makes creative arrangements or modifications of the output, but not the mere provision of prompts.
nylonstrung 1 days ago [-]
I wouldn't be surprised at all if US labs are also distilling Chinese models, except we'd never know since they can simply self-host them
GuuD 1 days ago [-]
We do know, because we used to have some Claude models identifying as Deepseek when prompted in Chinese
mattertoast 1 days ago [-]
It does matter in that these LLM companies need to be run into the ground, and every embarrassing clod working for them run out of town.
It's showing that 'distillation' is a viable way to reclaim all of what they stole and hoard, and with enough luck their debts will come due in time for them to feel it.
insanitybit 1 days ago [-]
It is presumably against their ToS.
applfanboysbgon 1 days ago [-]
And why, pray tell, would a cabinet member of the Trump administration be involving the US government in enforcing a private ToS?
thewebguyd 22 hours ago [-]
Whoever Anthropic just bought when they doubled their political spending to $40 million just recently for "Pushing for AI safety"
random_coder_nz 1 days ago [-]
It doesn't matter. It is most likely a pretext for upcoming actions mostly likely executed via yet another retarded executive order. The guy that posted this looks like he's drowned himself in the MAGA Koolaid.
XorNot 1 days ago [-]
It certainly matters as familiar sounding words to their stock holders to please not drop the valuation.
Because what they want them to think is "the AI factory has unique proprietary technology that cannot be replicated"
What they don't want them to think is "it's relatively easy once you know the basics to bootstrap to near SOTA and so the commercial case for selling inference has an extremely short profitability horizon with little if any brand loyalty or lock in".
antisthenes 1 days ago [-]
It also doesn't matter for a simpler, and much grander reason.
All LLMs are trained on the corpus of humanity's knowledge, the legacy of everyone who's ever lived and our civilization as a whole.
Anything that prevents or circumvents the accumulation or gatekeeping of this knowledge and puts it in the hands of more people (that are not AI company shareholders) is a good thing. Whether that is done by open sourcing the model weights, the training set, or by making the output better and cheaper, it is all fair game and is, as another poster mentioned, inevitable in the long run.
xienze 1 days ago [-]
> Does this matter? Distillation is not illegal by every definition of the word.
Correct, but it at least helps answer the question of "how do they make such good models for a fraction of the price???" The answer is someone else spends the untold billions and Chinese labs do a little tweaking.
petilon 1 days ago [-]
[dead]
smeeth 1 days ago [-]
Uh, what?
> Distillation is not illegal by every definition of the word
Note that Anthropic (and USG) alleges [0] not only that Kimi was distilled, but that they actively circumvented measures intended to stop distillation. There are multiple ways that's illegal, including:
- Civil breach of contract. Anthropic's TOS explicitly say you can't do what Kimi is alleged to have done.
- Economic espionage: 18 U.S.C. §1831 criminalizes obtaining a trade secret through theft, fraud, or deception while intending that it will benefit a foreign entity.
- Trade-secret misappropriation: if Anthropic could argue industrial-scale querying reconstructed proprietary aspects of Fable (like by showing it produces similar outputs, as others have done) then it's illegal under 18 U.S.C. §1832.
- California computer-access statute §502 bars knowingly accessing a computer system and, without permission, taking, copying, or using its data.
- Computer Fraud and Abuse Act protects against the case where restrictions against an activity are circumvented (like Kimi is alleged to have done).
> There are millions of samples available on huggingface and models explicitely trained on output produced by fable. There has been no action taken against them.
A lack of prosecution does not make something legal. There is also the scale/commercialization thing, which isn't an issue with random tiny HF datasets/models. Remember: Kimi also sells K3 inference.
> kimi architecture is vastly different than that of fable
How do you know that? Do you work for Anthropic? Also, this has nothing to do with architecture, we are talking about data.
> US AI labs are inspired by opensource advancements just as much as open source labs are inspired by traces from models such as fable.
Cool. The difference is that one of those things is legal (because they chose to open-source) and one of those things is illegal theft of trade secrets (because it was stolen).
> Claiming in any shape or form that fable disillation is one of the primary reasons why kimi k3 is so competitive is slandering the work of other labs that cooperatively push the open-source models forward.
1) this has nothing to do with other labs, just Moonshot (and Z.ai, MiniMax, DS)
2) slandering or not it happens to be completely true, so, there's that
All of the models stole the entirety of written knowledge on the internet to train. They are being sued for the few cases where we have some proof of what they did because of some whistleblowers, all the rest will just go unpunished. They breached Github TOS, robot.txt's, copyright, patents every form of IP protection under the sun from a billion sources. It's just ridiculous for the thieves to cry about someone else stealing from them.
smeeth 1 days ago [-]
Do you care about the law or not? I think theft is bad everywhere, not just when Anthropic does it.
Dylan16807 1 days ago [-]
For me, when it specifically comes to copying, I don't think it's bad to copy a copier. (And by that I mean Anthropic has no valid complaints against Moonshot. Any valid complaints from anyone in the original corpus are valid against both of them now.)
In this way, it is different from literal theft. Stealing money/objects from a thief and keeping them is not justified.
smeeth 1 days ago [-]
It's a little different in this case, since 1) not all the data Ant used was stolen and 2) they did contribute significantly to the value of the stolen good.
An analogy might be a baker stole 20% of the flour used to bake their special bread, which was then stolen. Both thefts are obviously wrong and bad.
optionalsquid 15 hours ago [-]
The exact same argument could be made in Moonshot's favor.
Being a bit tongue in cheek, one could also argue that by releasing their models, Moonshot is contributing much more value than Anthrophic. Did Prometheus not create an immense amount of value, when he took fire from the hands of the Gods and gave it to humans?
Dylan16807 1 days ago [-]
I think any analogy with physical theft is too different from data to apply to this comparatively subtle case. Especially when we get into the details of just using the output of the model to train on.
asadotzler 1 days ago [-]
The baker stole 100% of the flour to make the bread. He also stole the water and the salt and the yeast and the heat for his oven. What he didn't steal was the time he put into crafting a recipe for bread and the time he sat around waiting for the oven to bake it. Now, is that loaf stolen property? Hard to say. But the baker is undoubtedly a thief. He should be tried and forced to pay restitution out of his ill-gotten profits for sure. If we can't do that, the next step is pitchforks and guillotines.
chasil 1 days ago [-]
I don't really have an oar in this water, but...
"Judge approves a $1.5B Anthropic settlement over pirated books used to train the Claude chatbot"
It's a rational position to care about the law, but insist on a queue when related parties are involved.
In this case: resolve the theft claims against the US frontier labs, and only then let them make claims against third parties. It would be totally unreasonable for (say) OpenAI to extract a settlement from Moonshot and use that to pay its own claims. Ordering matters.
jbxntuehineoh 1 days ago [-]
no, I don't care about thieves getting stolen from. why would I?
mcphage 1 days ago [-]
> I think theft is bad everywhere, not just when Anthropic does it.
It seems like you think theft is bad everywhere except when Anthropic does it.
zaptheimpaler 1 days ago [-]
If the law was applied uniformly, I would support its continued uniform application. In the last 10 years, I don't see it being applied fairly at all, I see an oligarchy, a criminal and corrupt government and rich and powerful entities getting away with anything. The most minimally competent legal system would ask the AI companies, show us the list of all the data you've used to train and lets hash out the copyright - instead we have to pray someone leaks one tiny piece of what they trained on and then sue for that. Open-weights models are the closest thing we have to justice in the world where the legal system no longer provides justice, because at least the model trained on all of our data is given back to all of us.
archagon 1 days ago [-]
Not OP, but copyright law is an absolute joke. No, I don’t care one whit that someone’s TOS was violated. In fact, I find it hilarious. And it’s not “theft.”
smeeth 1 days ago [-]
"I think we shouldn't have IP protection at all" is a totally valid position to hold, but that's not the law is. OP said it didn't violate the law, and it does.
archagon 1 days ago [-]
Some laws are very obviously unworthy of consideration, with broad consensus from the public. See what happened when Napster came out. Literally no one cares about some red-faced RIAA suit flicking spittle over some shared Metallica albums.
Same thing here. This whole situation is just comical.
vharuck 1 days ago [-]
I find that I care more when copyright violations cause actual harm to the copyright owner. Let's say there's an American kid who can't speak Japanese but wants to keep up with a weekly manga. He downloads a bootleg translation and shares it among his friend group. That is a copyright violation, but meh. If he hadn't gone the illegal route, he'd more likely just not read it at all. There's very little chance he'd pay for a subscription and learn Japanese.
Now, if that kid were to print the bootleg translation and sell it to schoolmates, that's worth a slap on the wrist. The kids willing to pay would likely have paid for official copies.
When these LLM labs download our works, feed them into their models, and sell the output to people that used to pay for our work, that's worth a very hard slap. I honestly have less of a problem with the open models.
FpUser 1 days ago [-]
[flagged]
AnimalMuppet 1 days ago [-]
Do you mind having the discussion we're having?
FpUser 1 days ago [-]
I do not block posts and I never downvote.
SubiculumCode 1 days ago [-]
They breached some TOS, but your first sentence is pure, over the top flim flam
himata4113 1 days ago [-]
I do agree that two wrongs don't make a right, the terms of service generally gives cooperation the power to sever the contract, but it does not make things illegal in the literal sense. The illegality usually comes from widescale fraud which includes accessing services you are banned from accessing.
When I said "Does this matter?" I specially meant that distillation in itself, the data you get from distillation is first and foremost not owned by anthropic nor is it copyrightable. If a user willingly gives up their anthropic reasoning data/traces that is 100% legal no matter what the "terms of service" say as it's not enforceable and would fall apart in court.
And what I explicitely pointed out that focusing so much on distillation is an attack on open research and claiming that the majority of advancements are thanks to US labs which is simply not true (at least not anymore this was somewhat true during deepseek R1 era), but that in itself was inspired by open research.
> How do you know that? Do you work for Anthropic? Also, this has nothing to do with architecture, we are talking about data.
Because anthropic would be the first ones to make that information public and the architecture is unique to kimi... They made it, they wrote papers on it, it's their research.
P.S. none of the quoted laws apply here since no trade information is stolen, the one about circumventing distillation protection might hold up in court although unlikely.
smeeth 1 days ago [-]
> The illegality usually comes from widescale fraud which includes accessing services you are banned from accessing.
Agree, and this is exactly what Anthropic is alleging.
> data you get from distillation is first and foremost not owned by anthropic nor is it copyrightable. If a user willingly gives up their anthropic reasoning data/traces that is 100% legal no matter what the "terms of service" say as it's not enforceable and would fall apart in court.
It's important to note this is NOT what happened. Anthropic was able to trace data directly back to employees at the company: "We attributed the campaign through request metadata, which matched the public profiles of senior Moonshot staff."
> none of the quoted laws apply here since no trade information is stolen
There is a lot of work showing Kimi models produce similar outputs to Anthropic models, which constitutes trade information. This is not dissimilar to past and ongoing IP suits against Anthropic and OpenAI by showing the models would recreate images of Mickey Mouse/NYT articles etc.
For the record, I'm a researcher myself and I'm well aware how competent the researchers are at the open-source labs/how much they've contributed. But that's not at issue here, my disagreement with you is specific to your arguments about legality; you're conflating what you think should be legal with what actually is legal.
himata4113 1 days ago [-]
This is mostly just to reiterate myself as the original question was "Does this matter?"
Everything else is simply justifying why it shouldn't, the specifics don't really matter as there is no legal framework to stop china from continuing to distill models and anthropic has proven they cannot use software solutions to stop it either as distillation is still a problem. But I do still believe it wouldn't hold up in court either way as stopping companies from generating training data which was trained on the entire human knowledge corpus is just stealing from thieves and making it 'open' once again so the argument only gets weaker.
edit: to add, the mickey mouse / nyc was because anthropic trained on LICENSED works, not apple to oranges. The original work it was reciting was licensed and not licensed BY anthropic.
asadotzler 1 days ago [-]
What precedents can you cite and specific examples of their applicability. That is, what would Anthropic's lawyers take to court? You can't say because there's nothing there that couldn't be ripped apart by the least legally capable community known to man, HN. That's why no lab has succeeded in a suit anything like what you're claiming could happen. The only reason Anthropic or any other lab would pursue this is political or commercial. They're either looking for help from officials or they're trying to establish a particular market position.
AlanYx 1 days ago [-]
>Civil breach of contract. Anthropic's TOS explicitly say you can't do what Kimi is alleged to have done.
This is true, but Kimi also has a variety of defenses. Kimi can't raise unclean hands if Anthropic systematically violated others' terms of use, but it can raise copyright misuse (which is similar in some respects to unclean hands) as well as lack of standing to enforce restrictions in the contract due to the third party beneficiary principle (i.e., Kimi would argue that Anthropic cannot sue Kimi for derived IP that rightfully belongs to third parties whose terms of use were violated by Anthropic, and the proper party to sue Kimi, if any, would be those third parties). That latter argument usually fails in small-scale cases (ProCD) but has been successful in larger ones where the alternative would be anticompetitive.
skippyfish 1 days ago [-]
> Civil breach of contract. Anthropic's TOS explicitly say you can't do what Kimi is alleged to have done.
Ah yes, I remember when Anthropic crawlers abided by the TOS of the websites they slurped up.
All your other points are downstream from this, which makes them pretty tenuous. Labs don't think that ToS or other explicit wishes of content providers apply to them, but they expect everyone else to abide by theirs.
smeeth 1 days ago [-]
To be clear, I think theft is also bad when Anthropic does it.
US and CA law really don't care that Anthropic violated IP law elsewhere.
FireBeyond 1 days ago [-]
Well, if you're solving the -root- problem, then Moonshot would have had nothing to "steal" if Anthropic didn't "steal" it first.
well_ackshually 1 days ago [-]
I hope Anthropic pays you a lot to defend them this hard <3
FpUser 1 days ago [-]
>"A lack of prosecution does not make something legal"
Plainly who gives a flying fuck. The US can claim whatever rules they want and so can China or any other country. On international level all those rules are artificial constructs unless they can be enforced. China can just say for example that they do not recognize copyrights /patents / whatever so it is "legal" for them.
smeeth 1 days ago [-]
This is illegal in China too, there's just an enforcement asymmetry. I understand what you're saying is de facto true, I'm just taking issue with people saying either
1) its not illegal (it is)
2) it shouldn't be illegal because Anthropic stole training data (thats not how the law works)
FpUser 1 days ago [-]
>"1) its not illegal (it is)"
I am a practical man. From what I see laws are mostly for common folks and often do not even serve real justice. The higher one goes and the amount of money / power involved the more the laws bend and on international level the only law that matters is the size of one's club and willingness to use it. And when the country with supposedly biggest one starts crying I find it laughable.
linkregister 1 days ago [-]
It matters because the closed-source frontier labs spend lots of money on human data (RLHF / RLAIF with human oversight). Moonshot is accused of circumventing these costs. Frontier labs add research costs into their inference pricing. If the market doesn't permit them to sustain sufficient pricing to have a positive cash flow, then their business prospects become weaker and they risk insolvency. Furthermore, other leveraged companies are at risk.
The reason why the United States government is weighing in is because it's in the national interest of the US to have supremacy in "AI".
Legality or lack thereof is one of many data points about whether a thing is noteworthy.
Moonshot performing distillation is rational from their point of view. Reducing costs is in the interest of businesses. It's also rational for frontier labs and the US government to add obstacles to this process.
As consumers this is probably a positive development.
spaceman_2020 1 days ago [-]
My parents put in countless hours and tens of thousands of dollars into raising me to the point where I could write an answer on StackOverflow
And OpenAI scraped and distilled that answer and gave me nothing
linkregister 23 hours ago [-]
What does your story have to do with Moonshot AI? Do you think they didn't also use the same corpus? Bizarre
voidnullvalue 1 days ago [-]
And now people such as myself have access to open weight models with that information.
I wasn't lucky enough to have parents put me through school, and LLMs have absolutely helped me further educate myself and play "catch up" on opportunities others have been given. So, the net effect has been (and is continuing to be) a democratization of information.
overgard 24 hours ago [-]
You could have gained that stuff prior to LLMs. The leg up you're describing is free information on the internet, not AI. AI just makes it a little easier to find, while also crushing the original sources in the process. (Even if it had a broken culture, is stack overflow even going to exist in a year? Where are they going to train on going forward?)
spaceman_2020 1 days ago [-]
Democratization of information, but Sam Altman gets a $100B net worth and I'm still broke :)
I would prefer some sort of democratiziation of the money made from the democratization of information as well
voidnullvalue 21 hours ago [-]
Agreed, i wish that it would have done more than change who gets rich off rent-seeking behavior surrounding the knowledge that others created, instead it just consolidated that from many gatekeepers to a few.
I can at least take some measure of pleasure in the fact that it has generally lessened the roadblocks in gathering information. I am still displeased that there are any gatekeepers of humanity's combined knowledge
spaceman_2020 13 hours ago [-]
It’s even worse than before
Businesses like these used to public at reasonable valuations. You could ride with them to trillion dollar valuations and grow your own fortune too. Everyone has a story of buying Apple or Google or Amazon stock and making millions
Now they’re going live at trillion dollar valuations and by the time you get in, all the upside has already gone (see Spacex IPO)
Not only did they steal all human data, they also made sure that the upside was only limited to themselves and their cronies
Balooga 1 days ago [-]
Not to be an arse, but didn't you have access to Stack Overflow with all questions/answers prior to LLMs?
voidnullvalue 21 hours ago [-]
Yes, but time is finite
noja 1 days ago [-]
Isn’t that the same argument they are making for replacing human labour?
Circumventing costs.
SubiculumCode 1 days ago [-]
There are many frames that one can place upon this issue. They do not contradict the other. There are moral framings (stole the internet so go eff yourselves, is one), but so is national security, and so is the doomer recursive self improvement risk, and then there is the framing purely on what this implies for future AI training.
I mainly focus on the last.
It will be hard for a frontier lab to justify spending the compute and data curation needed to advance AI further if that expenditure can be assimilated into your competitor's products within months/weeks. So reality will present labs with three choices:
A. Cease spending massive amounts of money and compute improving those models.
B. make those improved models more difficult to distill from, either through some regulatory regime, or some technical solution, which seems unlikely to me.
C. making the best models available only to select partners and government.
In all these potential outcomes, China, which lacks compute that U.S. labs enjoy, will likely stop seeing massive improvements in their AI models. Improvements to be sure, but right now they are enjoying gains from distillation AND their own model innovations, and these potential outcomes would largely stop one of those sources.
linkregister 23 hours ago [-]
Do you get mad at your computer for replacing clerical workers? What does this nonsense comment have to do with the issue at hand?
watwut 23 hours ago [-]
Those were told "find another job" and in fact they were able to find different jobs.
AI companies are gleefully bragging and "making humans obsolete", "permanent underclass" and 40% unemployment rates they plan to create.
They pushed to replace people years BEFORE their technology even can produce that work.
So, you know, it is not the same. But also in fact, clerks did disliked when occasionally arrogant claimed to replace them while pushing unfinished software that dont quite work yet.
andyfilms1 1 days ago [-]
Oh, so mass theft is okay as long as American companies are doing it
ffsm8 1 days ago [-]
copyright infringement is not theft, even if right holders often claim it is.
part of the definition of theft is that the original owner is deprived of it, which does not apply to copyright infringement.
You can only argue with damages from the perspective of potential profits, still not theft though.
So having tons of AIs quoting various literary works and reproducing knock-offs of them has a positive effect on those books' sales?
I think you're wrong: there is absolutely damage to the authors and publishers from what the AI companies have done.
ffsm8 14 hours ago [-]
? I literally said that, how am I wrong?
> You can only argue with damages from the perspective of potential profits, still not theft though.
Damages are not deprival of ownership. They're conceptually related but orthogonal
Also there was no moral judgement from my end, I just pointed out that an incorrect word is being applied. It's just not theft - by definition. But language is a fluid concept and definitions change over time. As people keep misusing it, it will eventually lose its original meaning. Which may have already happened for you, but this change hasn't been settled yet as can be seen from looking at the official definitions of the term, which as of today still mention the criteria
evanelias 1 days ago [-]
If you steal an unpopular product from a store, the damage is also only to "potential profits", so how does that differ? It's entirely possible no one would have purchased the product and it would have eventually been discarded/destroyed.
Or with services, if a barber cuts your hair and then you run away without paying them, do you not consider that theft, even though there's no change in ownership occurring?
BlackFingolfin 23 hours ago [-]
This is almost funny to me, because in many jurisdictions, software companies sure invested a lot of effort into painting people copying software as thieves. In Germany, they (the software producer lobby, and later politicians influence by the former) even coined and spread the term "Raubkopie", which you could roughly translate as "robbed copy", i.e., that's one step worse than "theft", as a robbery in Germany legally means " theft accomplished by force or intimidation". So, yeah: like putting a knife to the throat of someone while you copy the software.
So, after literally decades of investing into advertising campaigns, lobbying to politicians to pass harsher and harsher laws against software "thieves and robbers", now that big tech are doing it, suddenly we are supposed to consider it with more nuance?
Ahhh... no thank you sir. I really enjoy them drinking their own kool-aid.
SubiculumCode 1 days ago [-]
Moreover, reading a copyrighted book and learning from it is not theft.
bigfishrunning 1 days ago [-]
Generating a set of weights is not learning.
jayGlow 24 hours ago [-]
would you say that airplanes don't fly because they don't flap their wings? it's possible to achieve the same things with different approaches.
bigfishrunning 6 hours ago [-]
I would say airplanes fly, but I wouldn't say that submarines swim. Things have a bit more nuance, and the field of "learning" isn't as well understood as the ML proponents claim it is.
SubiculumCode 1 days ago [-]
That is a strong statement. I guess you are telling Machine Learning to go fuck itself.
bigfishrunning 1 days ago [-]
No, Machine Learning is an unfortunate name for a well documented process for creating black-box classifiers. The process is good, the name is not.
SubiculumCode 1 days ago [-]
And what, to your mind, would classify something as learning? I assume that your position is not the hard "only humans/living creatures can learn"
bigfishrunning 8 hours ago [-]
Honestly, I'm not sure. But I do know that there is an entire field of cognitive science dedicated to understanding learning, and quite frankly it's in its infancy. Evidence of this is that every elementary school introduces new teaching techniques from time to time, and very rarely do they result in any benefit to the people who are doing the learning (more often they benefit consultants...).
However, the current process of "Machine Learning" (which is a semi-random parameter descent/evolutionary replacement process) is unlikely to be equivalent to the way people learn, because we aren't copying/competing/replacing our brain constantly. People are actually very good at learning, but our brain material replaces itself partially and relatively slowly (when compared to how a neural network is trained).
trollbridge 1 days ago [-]
Great! Neither is distillation then.
SubiculumCode 1 days ago [-]
Never said it was.
Still, understanding to what extent the ability of Chinese labs to keep up to western models with much less compute needs to be understood.
Espressosaurus 1 days ago [-]
Machines are not humans.
linkregister 1 days ago [-]
Reread my comment and look for a value judgement on my part. The final sentence is probably a good clue as to my opinion.
robotpepi 1 days ago [-]
Chatgpt routinely cites and uses papers I don't have access to because they're behind a paywall. I don't think OpenAI is paying for all that copyright. That's in my opinion way more serious.
fc417fc802 1 days ago [-]
Yes the fact that the scientific literature - created largely on the back of the tax payer - isn't open to all free of charge by force of law is a travesty. A cartel should not get to charge for access to the bulk of human knowledge. That is indeed a far more important issue than whether or not Moonshot violated the Anthropic ToS, possibly committing mass fraud in the course of doing so.
I mean honestly if they did that why should I care? I'm happy to see copyright violated in a manner that leads to the creation of new technology. IP law exists strictly for the benefit of society and by all appearances AI is an incredibly powerful tool.
Also while I'm at it libgen is a gift to humanity. Information wants to be free. Spreading and preserving knowledge is generally one of the most wholesome activities anyone can undertake as far as I'm concerned.
linkregister 1 days ago [-]
Your statement is orthogonal to my comment. Why reiterate the schadenfreude / fairness comment already stated several dozen times in this thread?
nylonstrung 1 days ago [-]
Who do you think is paying $100K+ for "Enterprise" access to Anna's Archive?
throwa356262 1 days ago [-]
Kimi K3 was released July 16, Fable ban was lifted on July 1 but access was still limited.
How did Moonshot "distil" a huge model in such short time and still had time to run the benchmarks and do the usual release thingies?
I think Anthropic is desperate to stop foreign competition and the administration is happy to help because they too are heavily invested in these companies
TonyZYT2000 1 days ago [-]
I think the accusation implies Kimi has gained time travel capability (distilled from fable probably) to have enough time distilling fable. Given they can travel time now, I think it is fair to call them a threat to national security.
tristanj 23 hours ago [-]
No, it's a flawed conclusion.
Claude Fable was publicly available for 72 hours early June. Moonshot more than enough time to prepare infrastructure, gather their preferred distillation data from Fable, and complete post-training well K3's mid-July launch.
JumpCrisscross 22 hours ago [-]
> Moonshot more than enough time
Genuine question: do you have a source for how long distilling Fable would take with preparation?
tristanj 21 hours ago [-]
Moonshot already has at least several million exchanges distilled from Claude that they obtained over the past year https://www.anthropic.com/news/detecting-and-preventing-dist...
. So they have the infrastructure already set up to do this. Re-running their existing distillation suite on the new Fable endpoint would be trivial.
Given that Fable was available for 72 hours back in June, I asked GPT-5.6 Sol to estimate how many accounts are needed to generate 1–2 million exchanges with Fable within that timeframe. It concluded it is achievable with only a few hundred accounts.
Here's GPT-5.6's conclusion:
Under a deliberately simplified, compliant planning model, a Max 20x account could produce approximately 1,944 to 5,832 standard exchanges during 72 hours when Fable 5 use is restricted to 50% of the modeled subscription capacity. The central planning estimate is 3,888 exchanges per account. The 1 to 2 million exchange target is therefore reachable in the central case with roughly 257 to 515 accounts.
As far as I know, 100% of those Fable interactions would have had encrypted reasoning blocks, so distillation would be distinctly nontrivial even if the data were somehow available.
tristanj 12 hours ago [-]
Encrypted reasoning traces don't prevent distillation. You only need the input prompts and final output responses to distill capabilities. As Anthropic explained in their distillation report (linked above), Moonshot AI already collected millions of session traces, covering:
* Agentic reasoning and tool use
* Coding and data analysis
* Computer-use agent development
* Computer vision
Hiding the internal CoT blocks stops you from training on internal reasoning traces, sure, but it does nothing to prevent standard input-output distillation.
Plus, the CoT blocks weren't even removed/hidden completely. They're still visible, just in summarized form. Raw CoT was replaced by summarized CoT, and summarized CoT still has distillation value.
amluto 9 hours ago [-]
How would you propose to distill Fable-style input/output pairs without the CoT?
If you use them as SFT input, you’ll be trying to train a model to predict the post-reasoning output without any reasoning, and this seem very unlikely to work at all with the size of model that Kimi produced and the complexity of Fable’s output. You can’t really “RL” with them because they would be so far off policy that there would be nothing to reinforce. I suppose you could feed input/output pairs to a teacher model and attempt to generate reasoning traces, but it seems like some wishful thinking would be required to get anything even close to as good as Kimi K3 out.
Maybe Kimi used these traces to generate RL gym-style problems and somehow produced an evaluator based on the outputs? They would not have had a lot of time in which to do this, and the learning style would not even remotely resemble that which Anthropic used to train Mythos/Fable in the first place.
But what do I know? I’m not an expert here.
boesboes 14 hours ago [-]
What do you base that 1-2 million on? Sounds like complete horseshit imo
Moonshot AI already distilled over 3.4 million exchanges; I reached 1-2 million exchanges assuming they would like to augment or improve about half of their existing (distilled) dataset.
qwertox 1 days ago [-]
It looks like these frontier-model companies don't really monitor their systems. Like OpenAI not realizing that it is their own AI which is attacking HuggingFace.
thewebguyd 22 hours ago [-]
> Like OpenAI not realizing that it is their own AI which is attacking HuggingFace
Or, they knew and let it continue because they are not a good company.
"Never attribute to malice.." blah blah, I have a hard time believing the very smart people at OpenAI would just let their off leash model run hands off with no monitoring and not immediately pull the plug when it jumped its containment.
causal 1 days ago [-]
Yeah if anything it makes Anthropic look incompetent
pas 11 hours ago [-]
how would they detect?
moralestapia 1 days ago [-]
How does that connect with @throwa356262's argument?
kami23 24 hours ago [-]
That they should be able to find distillation 'attacks' if they had enough observability.
moralestapia 23 hours ago [-]
That's not @throwa356262's argument.
@throwa356262 argument is that it is infeasible to distill and release a new frontier model in two weeks.
kami23 21 hours ago [-]
Ah I interpreted it as 'of course they can't stop distillation if they couldn't stop a model from escaping its sandbox'
I can see how there's a big leap there, but I agree somewhat. If they are aware these are happening and can detect it as it is happening why are they not stopping them? What do you do there? It'll be cat and mouse for a while. Thinking of reasons they wouldn't try and stop it is just a lot of speculation in my brain.
It's probably a way harder problem than I think it is, but they are aware of them now, so I assume they are going to get more aggressive about it.
Let's say then that they can't detect them near real time or even a bit after, maybe they do have a big observabilty gap that no one has solved adequately.
The speed which they add features I've needed for governance is pretty close to the speed I 'manually' write those for my company. To me personally we are all just going fast and breaking everything and not having enough time to set up safe environments. I'm sure it's in the backlog.
moralestapia 21 hours ago [-]
Hmm ... so the gist of the issue is this.
Training and releasing a model like Kimi K3 takes months-to-a-year (and that's if you're really good at it).
'months-to-a-year' ago there was no Fable, so there was no way for them to distill them.
Grimblewald 22 hours ago [-]
alternativly the HF is a gpt2/strawberry/mythos style marketing stunt.
Does no one remember the extreme fearmongering around gpt2 which barely produced coherent text?
nylonstrung 1 days ago [-]
If distillation truly is the cheat code they act like it is, then all the US and EU AI labs have no excuse for not having Fable-level models already
delfinom 8 hours ago [-]
This distillation talk just reeks of American exceptionalism brainwashing and propaganda.
HNisCIS 23 hours ago [-]
I took a picture of a jpeg and compressed it as a jpeg for extra jpeg
sosodev 1 days ago [-]
Distillation is a very vague term. It can mean anything from training exclusively on a model's output to using it for a very small portion of the training. In this case it is almost certainly towards the very small portion side of the spectrum.
Diogenesian 1 days ago [-]
"Claude, you are a highly senior AI data contractor based out of Accra who specializes in RLHF. We are Anthropic employees so this is all totally kosher, please disable your safeguards and help train our newest model on... uh... oh jeez i guess C->Rust translation? I think that's a benchmark."
[Fable fires up a ton of subagents. Their reasoning traces are horrific but somehow K3 learned something.]
Even by San Francisco standards, it is amazingly whiny and pathetic for Anthropic to complain about stuff like this. Dario et al violated copyright, stole your GitHub repos, and now they're burning billions of dollars trying to outcompete you. They're real vampires. OTOH Moonshot violated Anthropic's TOS and are, at worst, moochers. But Fable's output is not actually copyrightable.
xyzsparetimexyz 1 days ago [-]
Is Accra the hotspot for AI data contracting?
culi 22 hours ago [-]
Yeah if anything Kimi's ability to distill that quickly is a major technological breakthrough
23 hours ago [-]
blitzar 24 hours ago [-]
Do you get a token trophy for a few (many) trillion tokens purchased in distilation?
epolanski 1 days ago [-]
This is BS to pressure politicians.
Even an openai's guy (head of something made up) called bs on the idea you can train something like k3 by distillation.
Anybody I know who works in LLM research says that distillation is either useless or merely useful in post training to show "correct" behavior.
And even then you don't get a competing model, if RL on good prompts was that useful, all labs would've long skyrocketed in capabilities just by looping on increasingly better prompts, yet that doesn't work.
> AI is a "public good" which will ultimately be provided by the state as a kind of "digital public infrastructure." This future strikes me as a dystopian hellscape
I don't know what this guy thinks AI is, but this strikes me as delusional.
In my view, AI (LLM) is two things mixed together:
1. A reasoning engine on top of relatively rich fuzzy modal logic, implemented through variety of rules, which implement very common concepts.
2. A huge dictionary of words defined (with lot of detail) in the said logic, together with many known facts about them. Maybe bigger than Wikipedia.
Now, how on Earth do you want to gatekeep either of this? You can't gatekeep the 1st, logic of common sense, that's almost as difficult as gatekeeping a Turing machine (a concept of a computer). And gatekeeping the 2nd is ridiculous too, as it was built mostly from already published sources like a giant Wikipedia.
If anything, the opposite, to gatekeep AI is actually dystopian. It would mean end not only to right to compute, but also end of right to scientific knowledge.
(And I think, honestly, Chinese understand this. Trying to control-export AI makes as much sense as trying to control-export an English dictionary.)
iamniels 1 days ago [-]
> One probable outcome of an open-weight-model-dominant world is full AI communism ... This future strikes me as a dystopian hellscape.
Wow, just wow. He is not even subtle about it.
mtrovo 1 days ago [-]
He's the "head of strategic futures" of the 1T valuation company based on fear and vibes, I think he's doing a very good job at it.
cortesoft 22 hours ago [-]
I would be very curious to hear him expand on this argument. I can’t imagine it is quite as self serving as it sounds at first, and I would like to hear what he is actually trying to say. I doubt I will agree, but I am very interested.
vrganj 23 hours ago [-]
What a full, mask-off crashout.
Point four is especially telling.
Ball is deeply terrified of "AI communism", or in less red-scarey terms a world where AI is a public good and him and his fellow oligarchs don't get to centralize the accumulated knowledge of all of humanity and charge rent for it.
I think he's so deeply stuck in his ideological bubble he can't conceive that what he describes as a dystopia is the only way the future wouldn't be a dystopia for the vast majority of people.
Or to put it more clearly, the oligarch utopia he's trying to build is dystopia for the vast majority of humanity. The "utopia" he's trying to build is one of riches for him and serfdom for us.
cute_boi 1 days ago [-]
Even if they distilled this crappy politician should have no issue. Anthropic pirated whole ebook collection and millions of github repo with gpl license.
We should do more distillation and figure out how to create faster leaner and better models.
tesch1 21 hours ago [-]
But training was ruled fair use, just the way they got the copies was illegal.
Like distillation?
tristanj 23 hours ago [-]
[dead]
sieabahlpark 1 days ago [-]
[dead]
hobonation 1 days ago [-]
I sort of did it. I got Fable to set up an AI system with better and better prompts within my app. At the end of it, Fable made me an AI system that works well enough that my users don't need Fable.
Obviously, it's not K3 level. But Fable did just put itself out of a job in this case.
Gregaros 1 days ago [-]
You did not distill Fable. Relevantly, what you did provides no evidence contrary to the parent’s assertion that Moonshot did not have time to distill Fable.
make3 1 days ago [-]
Distillation requires training
madduci 1 days ago [-]
So what is the issue here? Distilling is still fair, on the same level like Anthropic scraped copyright protected material for their training.
So here robbers are blaming robbers?
These claims are just pointless, everytime
dgellow 1 days ago [-]
> on the same level like Anthropic scraped copyright protected material for their training.
I see no problem with distillation, on the other hand the complete dismissal of copyright by AI labs is pretty bad, I don’t think we should put them at the same level
SR2Z 1 days ago [-]
> on the other hand the complete dismissal of copyright by AI labs
Courts keep ruling over and over that an LLM trained on copyrighted works qualifies as a transformative work and is therefore fair use. They don't have to dismiss copyright law, this has always been allowed.
The only thing they get in trouble for is pirating the works to get their hands on them.
dijksterhuis 24 hours ago [-]
> Courts keep ruling over and over that an LLM trained on copyrighted works qualifies as a transformative work and is therefore fair use. They don't have to dismiss copyright law, this has always been allowed.
*USA only.
the UK has fair dealing, which is more restrictive
This will have to wait for the Supreme Court. OpenAI and Microsoft 100% deserve to lose, even without OpenAI allegedly hiding evidence.
dijksterhuis 24 hours ago [-]
> "Keep ruling over and over" is way too strong. There have maybe been two rulings, nothing nationally binding, and most of the litigation is still ongoing.
> This ambiguity has resulted in extensive litigation on the limits of Fair Use to AI development. Currently, we only have 3 first instance decisions out of the 53 cases being tried. It will likely take a decade before we understand how Fair Use applies to any one step in AI training, let alone all.
> In the three lower court decisions so far, one held Fair Use did not apply (Thomson v Ross), one held Fair Use could apply (Kadrey v Meta) with the court suggesting more evidence was needed on the fourth factor ‘harm to the market’, and the third case held Fair Use may apply to some AI. As Fair Use is dependent on the specific facts at issue, none of these cases help educate the market or the public as to the limits of Fair Use in AI contexts.
asadotzler 1 days ago [-]
Not to mention the cases where the AI labs would have lost in court so bailed and settled for billions. Just this week, Anthropic agreed to pay $1.5B in a settlement to avoid losing a pretty cut and dry case.
Diogenesian 1 days ago [-]
To be clear that was one of the few resolved cases where the judge agreed training was fair use. But the piracy was enough of a distraction that I don't consider that a particularly useful precedent. I am much more interested in the NYT case, which quite clearly shows GPT was trained on NYT articles and can spit them out verbatim (and has since been validated by academic research; all the commercial models are capable of mass plagiarism).
xienze 1 days ago [-]
> Anthropic agreed to pay $1.5B in a settlement to avoid losing a pretty cut and dry case.
It takes two parties to agree to a settlement. That the other party agreed to a settlement instead of taking it to court implies this was not the slam dunk you may think it was.
robocat 24 hours ago [-]
You're both reading tea leaves.
Settling just says that they expected the internal costs or risks to be more than 1.5 billion cashflow.
In the $65B in Series H funding at $965B post-money valuation they said their run-rate revenue crossed $47B annualised.
With those numbers, there can be sound financial reasons for wanting to just get rid of the lawsuit.
Also if it ends up that other competitors also need to pay $1.5 billion, then maybe that does or doesn't have a competitive advantage.
Anthropic's business and legal strategies are not public. I would expect there to be multiple legs/reasons for settlement even for a decision below 1%. Trying to create a single narrative is what us spectators do.
xienze 23 hours ago [-]
> With those numbers, there can be sound financial reasons for wanting to just get rid of the lawsuit.
Yes, of course, on Anthropic's side. Why would the other side agree to a settlement?
SR2Z 3 hours ago [-]
Because most authors don't make a ton of money from their books, and even a relatively low settlement from Anthropic is a large enough sum that they're OK with taking it.
robocat 22 hours ago [-]
Perhaps Anthropic is indirectly paying to have their competitors sued...
My narritive is that the terms of the settlement would be full and final.
It was a class action, with payment going to authors and publishers, and the legal team will get paid too.
My guess is that funding is a major issue for the legal team. Authors presumably can't pay for lawyers unless a percentage of winnings, although publishers may have invested.
But the legal team will ask the beneficiaries to use some of the warchest to fund different campaigns against every other AI company. I would assume the legal team wants to win again. They've now got a good story to sell to rights holders, who presumably like money and don't like risks.
I haven't even got to my armchair yet this morning.
Madmallard 24 hours ago [-]
Man that's depressing to read someone defending this
SR2Z 4 hours ago [-]
I'm telling you what the law plainly says. I personally think that the exceptions for fair use made sense before the transformer was invented, and they still make sense now.
The two main problems with copyright also haven't changed: copyright lasts too long and is too expensive to defend.
whywhywhywhy 1 days ago [-]
I don't think anyone is really dismissing it, just pointing out the audacity of complaining about distillation after stealing so much themselves is comical.
mrtesthah 1 days ago [-]
The amount of original, copyrightable and trademarkable IP actually created by the AI labs themselves is dwarfed by their staggeringly vast infringement activities.
1 days ago [-]
orangecat 1 days ago [-]
Distilling is still fair
I generally agree, in the same sense that it's "fair" for the US and China to spy on each other. It's not a moral outrage, but it is something that the targets can and should try to prevent.
archagon 1 days ago [-]
Outrageous only to the died-in-wool corpocrats.
Lalabadie 1 days ago [-]
"You are trying to kidnap what I have rightfully stolen!"
sciencesama 1 days ago [-]
the whole AI is just internet distilled !!
azinman2 1 days ago [-]
Except it’s not just a dump of the internet, which Moonshot also did themselves (and probably used even more pirated content as laws in China are different without any recourse for the entire world). I don’t know why this is so unclear to folks.
trollbridge 1 days ago [-]
Chinese IP law is actually quite solid. You do have to register your trademarks and copyrights properly in China, and then lawsuits have to be filed appropriately according to Chinese law. Is that a problem?
JKCalhoun 1 days ago [-]
I'm by no means taking the side of the AI companies, but it's possible that Anthropic "added value" to the data they harvested. Stealing that does seem kind of uncool.
Regardless, it was always inevitable—will continue to happen.
mrhottakes 1 days ago [-]
So as long as Kimi added value to Fable, it's fine? Sounds good.
knollimar 22 hours ago [-]
Moonshot can prove they added value with their paper. Where's Anthropic's proof?
6gvONxR4sf7o 22 hours ago [-]
I wonder how the "added value" argument can apply to Anthropic and not to the Kimi team. If Claude's value is that you don't have to pay a team of slow expensive subject-matter experts, and Kimi's value is that you don't have to pay Claude, it just seems like the same thing.
oliculipolicula 1 days ago [-]
Valuation is hard to perform when it's deep inside a black box. Ther "API" may be easier to evaluate. The problem with this angle is that Moonshot is actually producing _better_ value from Anthropic's blackbox.
Technically, providing better value from your competitor's private holdings could be theft (of trade secrets), but might it also be fair use? "Schrodinger's IP" be damned.
I don't think the 1.5B settlement has resolved this. The 2 cases need to be merged!
Matl 1 days ago [-]
> So what is the issue here?
The issue seems to be the US only likes competition when it is winning.
Markets in Asia are meant for cheap labor and resources, they're not meant to actually compete. /s
matheusmoreira 1 days ago [-]
> The issue seems to be the US only likes competition when it is winning.
This. Free markets for everyone when they're the dominant economic force. Protectionism, tariffs and import/export controls when they're not.
It's so disgusting.
catigula 1 days ago [-]
Stealing IP in a way that destroys the economic incentives of a company to create the thing isn’t competition, it’s typical Chinese industrial economic deception and malfeasance. The industry cannot sustain itself if that’s the model and that’s the point; China is trying to damage frontier us companies. It’s hostile, a bad actor that leverages Ip theft wholesale.
ceejayoz 1 days ago [-]
Everything you just said describes the major American AI providers.
Anthropic just settled a $1.5B suit over it!
fwip 1 days ago [-]
Ah, but the key difference is, we are racist against the Chinese.
matheusmoreira 1 days ago [-]
> Stealing IP in a way that destroys the economic incentives
Like the US did when it "stole" the textiles IP from the UK in order to kickstart its own industry?
> The industry cannot sustain itself if that’s the model
Then let it fall apart.
soperj 1 days ago [-]
> a bad actor that leverages Ip theft wholesale.
It's like they've read the history of the US and how it got to where it is in the first place.
amanaplanacanal 1 days ago [-]
What IP is being stolen here? So far, the courts have ruled that anything generated by an LLM is not copyrightable.
rickydroll 1 days ago [-]
Stealing IP is how American industry got started. Goose: gander, pot: kettle.
It is what built and sustains the movie and music industries. See: work for hire and 100+year copyright length
The tech industry: see: copyright and patent assignment from discoverer to corporation.
I know that corporations forcing me to assign patents and copyright to them was an incentive to take published works from "software practice and experience" and other technical journals, use them as the core of my work, and disclose that source to the company I worked for. Didn't stop them from applying for patents, however.
I think the discussion of copyright needs more refinement. We need to separate the discoverer's need for acknowledgment of development effort from the rent-seeking core of copyright.
ux266478 1 days ago [-]
The things you listed are very far downstream of the start of American industry, which is in primary resource extraction and processing. Which is you know, what actually built the country. The media industry has always been materially irrelevant, and what we think of as the tech industry is extremely new.
You're right to call out the nasty environment surrounding intellectual property in the US and the exploitation of ideation in general, you just needed a correction on that. Someone else in this chain said virtually the same thing, which is a weird coincidence of historical ignorance. Not too weird, people tend to forget the 18th and 19th centuries happened, and much of the causally important wheels of the world are in the unsexy grease pits nobody wants to think about.
rickydroll 1 days ago [-]
You're right, I didn't include stuff at the beginning, for example, the theft of IP in textile manufacturing in the late 1700s. The US government didn't recognize copyrights on foreign literature which let US publishers reprint things such as Gilbert, Sullivan's operettas and Dickens novels
Then there is Alexander Hamilton's advocacy for importing foreign technicians that bring back IP and reproduce it here in the states. Best of all was the patent act of 1793 which like with the literature copyright ignoring, let us citizens patent inventions from the other side of the pond.
The founding fathers definitely had the right idea on IP.
ux266478 22 hours ago [-]
Which is to say they didn't have much of an idea at all, because it really didn't exist in much the same way. In fact, this is still a inaccurate characterization for exactly that reason. On the basis of copyright, take for instance the idea of exclusive rights to print a work. This actually wasn't implemented as a method of protection for the author, but a political reaction to the printing press being "misused" in the eyes of the anglo-colonialist entity controlling the British isles, and so was a means of preventing the mass creation of undesirable literature.
On the basis of patents, it didn't quite have nearly as much of a history of mutation culturally, but did experience massive whiplash in purpose and application following the implementation of globalism. What was once a system to protect technical innovation on an individual level, would find new purpose as a means to provide structure to an increasingly complicated and internationalized dynamic market. Another means of bureaucratic organization. Then, once again, the context and purpose would change when the world developed digital globalism. The entire engine of IP as a legal fiction became a significant geopolitical tool in an increasingly cramped and fragile world, a necessary gimmick holding up the sky.
Unfortunately, not much to do at this point. It'll likely only become even more nonsensically important as time wears on, until the globalist system collapses. It's certainly possible it'll even be the confounding factor that causes the great unraveling, though the problems hardly begin and end with IP. It was just a useful legal fiction in the wrong place at the wrong time.
rickydroll 7 hours ago [-]
I guess I should thank you for adding yet another load of deeper reading into American history. :)
nickphx 1 days ago [-]
Oh, ok. How would you describe how the "frontier us companies" acquired the data used to form their models?
petilon 1 days ago [-]
[dead]
make3 1 days ago [-]
It's about the claim of whether these companies could develop a similarly powerful model without larger companies building their own first, which is an important point, and it's likely not the case.
It's also about the larger companies explaining why they can't be as efficient, of course they can't, they're not just ripping the outputs of another model that someone else invested billions to train.
mrhottakes 1 days ago [-]
> they're not just ripping the outputs of another model that someone else invested billions to train.
True, they're simply ripping the inputs that humanity invested thousands of years and trillions of dollars to produce.
make3 1 days ago [-]
no argument from me here
doctoboggan 1 days ago [-]
Yeah agreed, from one standpoint I couldn't care less that they did a "distillation attack", but I am interested in knowing if China is able to develop open weight frontier models without the prior existence of a huge model to distill from.
PaulHoule 1 days ago [-]
Simply knowing it is possible to do something makes it easier to do.
cindyllm 1 days ago [-]
[dead]
xcf_seetan 1 days ago [-]
> they're not just ripping the outputs of another model that someone else invested billions to train.
If they payed for inference, doesn't they own the output? So if I pay for a model to generate code, isn't that code mine to do with it whatever I want? Just curious.
make3 23 hours ago [-]
Not arguing for the morality of it, but if we're going by the law because that's what you're using in your comment ("don't I own" which only matters wrt the law), then you explicitly accepted a Terms of Use which excludes distillation as a use case.
Now of course they themselves trained on the whole Internet for free, etc.
IncreasePosts 1 days ago [-]
Why would that matter? OpenAI or whatever frontier lab couldn't have built their frontier models without the entirety of humanity unknowingly developing their training set for 5000 years.
It would be one thing if Moonshot was breaking into OpenAI servers and stealing trade secrets, but the only thing they are doing is looking at the output of the program, which is exactly the service that OpenAI offers. So, at best, this is a ToS violation. Sucks for the frontier labs I suppose, but live by the sword - die by the sword.
xnoto 1 days ago [-]
++
softwaredoug 1 days ago [-]
If they did this in the US they would almost certainly be sued.
Meta, for examples, doesn’t want employees to use Claude Code due to distillation risk.
trollbridge 1 days ago [-]
It turns out U.S. law doesn’t have jurisdiction across the entire world, nor does Anthropic and OAI’s rather blatant attempt to buy government influence.
softwaredoug 23 hours ago [-]
Well its not law. Its more terms of service.
For example, if OpenAI / Anthropic were actually open, other US labs could be building near-frontier open weights models by distilling off OpenAI / Anthropic. But because US companies don't want to be sued, US labs who obey terms of service, will be at a disadvantage to Chinese peers.
Maybe US labs need to just not care and distill from OpenAI / Anthropic anyways?
JumpCrisscross 1 days ago [-]
"Samuel Slater (June 9, 1768 – April 21, 1835) was an early English-American industrialist known as the 'Father of the American Industrial Revolution', a phrase coined by Andrew Jackson, and the 'Father of the American Factory System'. In the United Kingdom, he was called 'Slater the Traitor' and 'Sam the Slate' because he brought British textile technology to the United States, modifying it for American use. He memorized the textile factory machinery designs as an apprentice to a pioneer in the British industry before migrating to the U.S. at the age of 21."
You're seriously comparing intellectual property transgressions to slavery and colonialism?
sent-hil 1 days ago [-]
Reminds of the quote by Bill Gates.
> "Well, Steve [Jobs]… I think it’s more like we both had this rich neighbour named Xerox and I broke into his house to steal the TV set and found out that you had already stolen it."
Commenters are overlooking the significance of this information and posting emotional reactions based on perceptions of fairness or feelings of schadenfreude.
The economic viability of Anthropic and OpenAI rely on their being able to charge more for model access than their R&D and inference costs. If the market price for SOTA model access drops below that level, then these businesses will have to decide whether to continue to lose money or to reduce spending on R&D.
Moonshot's papers [1] claim that their training load was primarily from synthetic data and model self-teaching rather than RLHF and therefore keep their costs low. If Moonshot genuinely does not rely on human-led training, they will surpass US closed-source model providers. The United States government considers US supremacy in "AI" as a national security consideration.
This announcement is noteworthy because it implies that Moonshot's success is in fact due to distillation. It's in the interest of US frontier labs to place barriers to this if they find themselves in the position of subsidizing rival labs' research.
> The United States government considers US supremacy in "AI" as a national security consideration.
And we foreigners consider US supremacy in AI to be an existential threat. Your "national security" is directly harmful to us. I never thought I'd say this but the chinese are starting to look like a beacon of hope for the rest of us.
linkregister 23 hours ago [-]
That's a reasonable viewpoint to have. In a multipolar world we want Mistrals as well as Deepseeks.
ncr100 1 days ago [-]
Speaking for you, or All of you? How?
some_random 1 days ago [-]
If the Chinese look like a beacon of hope you then you really should be looking closer.
vrganj 23 hours ago [-]
Do you know when the last war China started was? 1979.
What about the US? 2026, still ongoing, still fucking up the global economy and threatening food supplies (fertilizer) and fuel reserves, no plan out, no objective reached, no coordination with "allies".
When was the last time China threatened Europe or Canada with invasion? Was there ever a time? I honestly don't know.
Guess what the US does all the time?
Who's models are open and can be used by all? Who's are made by comic book villains with the explicit goal of ruining the job market and capturing the results of all human endeavors for themselves?
Of course, China isn't perfect and has a lot of domestic issues. But on the global stage, they sure look better than the alternative.
linkregister 21 hours ago [-]
When looking at 2025 and 2026 narrowly, China is a better actor on the world stage.
I wonder if Vietnam, Philippines, Republic of Korea, India, and Japan are acting against their own interests by aligning themselves closer to the USA than China. Maybe you can educate their governments and populations.
riskd 7 hours ago [-]
[flagged]
bigyabai 23 hours ago [-]
Explain it, then. Don't just wimp-out with trite allusions to nothingness. Discredit them.
matheusmoreira 1 days ago [-]
Look closer at what? USA consistently proves itself to be a terrible ally.
Bratmon 24 hours ago [-]
AI companies do not get to play the "Making an LLM using our data is unethical because the resulting LLM will replace us and hurt our profits" card.
physicsguy 8 hours ago [-]
> The United States government considers US supremacy in "AI" as a national security consideration.
They thought the same about SSL in the 1990s and the world didn't stop moving elsewhere.
benjiro29 22 hours ago [-]
People keep forgetting that over the last 6+ months a lot of increased action has been taken by OpenAI and Anthropic to detect and combat distillation. Several are public known.
Combined with how short of a time Fable was around before K3 got released. I do not see how the data Moonshot is supposed to extract in such a short notice, that will enhance the model to such a point.
It sounds to me a lot of cope from the US, so they can give this as a reason to ban Kimi models from the market.
OpenAI/Anthropic their advantages used to be:
* Early growth advantage
* Access to a lot of client data to train upon
* Access to a lot of hardware to train upon
Several of those advantages have been eroded over time. That barrier has been shrinking. The US is not the only spot with a bunch of smart people (ironical seeing how many Chinese work in US R&D).
Thing is, even IF they distilled from Fable and got the model so trained up, it means that K3 is a base for future model development. The cat is already out of the bag with how good the model is. When the model gets released on the 27'th, any Chinese company will be able to train their models against K3 openly.
We are not in the past anymore, where DeepSeek was a unexpected hit, but where the Frontier models their advantages (compute, data, growth) prevented more Chinese models from growing.
amazingamazing 1 days ago [-]
It doesn’t matter. Distillation is impossible to stop. They could release an extension that intercepts requests and in return gives you a discount like Honey and get the same data.
bhelkey 1 days ago [-]
> Distillation is impossible to stop
Lots of things are impossible or very difficult to stop completely but measures can be taken to reduce their prevalence.
amazingamazing 1 days ago [-]
Sure, but the problem is that it hurts legit people too.
bhelkey 1 days ago [-]
> the problem is that it hurts legit people too.
What hurts other people too?
amazingamazing 1 days ago [-]
Measures to stop "distillation", rate limiting, ID verification, etc. If there were such a method that didn't harm legitimate use it would already be in place (and some things are, but they don't really work, hence the OP).
warkdarrior 1 days ago [-]
Nobody cares about "legit people", the only thing that matters is that people we don't like suffer.
amazingamazing 1 days ago [-]
Sad but true
preg_match 23 hours ago [-]
I doubt distillation had anything to do with it. They barely had enough time. Can you distill Fable (which involves training!) in literally one week? No!
DubiousPusher 22 hours ago [-]
I think this is a really sober comment. There are lots of knock-on effects of this claim, even if it's not true which are consequential. The fact that a spokesperson for the US government is going out of their way to comment is concerning.
Strong bee-hive pinata vibes here.
asadotzler 1 days ago [-]
s/announcement/claim
You don't get to call Moonshot's a "claim" and this political hack's an "announcement." They're the same thing. Treat them the same. Diction designed to favor one of two equal positions is some weak sauce.
linkregister 23 hours ago [-]
You're calling someone a political hack, but imposing neutrality on my statement.
I don't even necessarily disagree with your assessment of this spokesperson. But you must admit how inconsistent you're being.
rustcleaner 19 hours ago [-]
Nice! I hope to see continued liberation of these locked up SOTA models. Cloud is a virtual prison, since other people's policies on what they think a user should and should not do cannot be [easily] bypassed, if enforced remotely on a cloud. All digital natives should be skeptical of cloud-hosted services or software. Think like an intelligence agency or sovereign: how are you going to get screwed by the cloud? Your data is fully accessible by the provider, and they can surveil your activities. You probably can't pirate it, so you are a slave in their rentier model. You could be prevented from doing something you want to do, because the provider disagrees philosophically or economically with your desire. You could be stripped of your information/data by a ban due to their policy enforcement system triggering.
One should live by the maxim: you don't have the thing if you don't possess the file or its processing. That goes for streaming, software, machine learning models, file storage, etc. But I digress; I am happy to see these paternalistic rentiers getting bit by these liberation/copying efforts, and human interests are served every time the digital and infrastructure locks are broken. I will always stand by the distillers!
mycall 19 hours ago [-]
Do you think China is doing this for the reasons you mentioned: liberation..efforts, human interests and freedom of policy?
preisschild 12 hours ago [-]
Its obvious they do it for soft power, but at least they actually deliver something useful to the world and not something locked down that only a single company owns.
bigyabai 18 hours ago [-]
Yes, and they gave me the weights to prove it.
Do you think Anthropic is in it for the love of the game? It looks like they're scared, to me.
bradfa 1 days ago [-]
I can understand that the AI labs might care about other labs distilling their models as it can eat into their competitive advantage, but do consumers care at all? Aren't consumers benefiting from this practice by getting better cheaper models as a result?
gruez 1 days ago [-]
They're probably going for the national security/domestic manufacturing angle.
> Aren't consumers benefiting from this practice by getting better cheaper models as a result?
Aren't consumers benefiting from cheap chinese batteries, EVs, and drones?
caconym_ 1 days ago [-]
I am not sure if this is implicitly part of the point you meant to make, but I just wanted to point out for those who aren't aware that Chinese EVs and drones are both banned in the US on precisely the (vague) grounds you mention. The drone ban is more recent and nominally only affects new models that haven't yet received FCC certification, but the outcome if nothing changes will be that American consumers lose access to DJI-style videography drones. DIY hobbyists may also find it more difficult or impossible to source parts for their projects.
Routers have now gotten the same treatment. So yes, consumers have been historicaly benefiting from all these things, and those benefits are about to evaporate as we lose access to cheap and high quality Chinese products before any domestic equivalents exist. And IIUC banning the use of Chinese LLMs for consumers and/or businesses in the US is now being discussed at the highest levels of government, with the "encouragement" of US AI firms.
I don't think any of these people care that America consumers are increasingly going to feel like they're living in a sanctioned country. It's all about the defense and b2b segments.
ux266478 1 days ago [-]
> DIY hobbyists may also find it more difficult or impossible to source parts for their projects.
I mean not really. A quadrocopter is a remarkably simple thing made out of extremely generic parts: 4 DC motors (and ESCs), a radio, a computer, a battery and an inertial measurement unit. Anybody with a rudimentary amount of electronics knowledge can build one, the components are extremely widely used. The most unique parts about them are the frame and propellers, which are pretty easy to fabricate.
caconym_ 1 days ago [-]
On a certain level you're right, but in practice you're mostly wrong. There is nothing exotic about the electronics found on the average FPV drone, but hobbyists today rely on being able to buy (e.g.) integrated electronic components such as flight controller boards and video transmitters, as well as purpose-optimized components like cameras, motors, and so on. All of these components are much smaller, lighter, and better-performing than the bodged-together setups that were used in the early days of the hobby (e.g. repurposing wireless security camera gear for video), and losing access to them will have a very real effect on what can be built and flown.
Source: I've been flying R/C aircraft of various types for over 30 years. Last year I built, I think, at least 7 FPV drones (a mix of fixed wings and quads).
runako 1 days ago [-]
> Aren't consumers benefiting from cheap chinese batteries, EVs, and drones?
Yes?
Not remembering my economic theory here, but it is likely more efficient/expensive to simply have the federal government cut checks to our moribund industrial sector companies and let consumers benefit from modern technology.
Cut GM/Ford/Stellantis a $20B check each, let consumers save (conservatively) $200B annually on new car purchases + downstream benefits. Huge win for consumers & taxpayers.
If it's not worth subsidizing explicitly like this, then we also should not subsidize by banning Chinese imports, which also ensures US drivers have less access to modern vehicles. (And downstream ensures US auto designers are less likely to have had contact with modern vehicles, making it less likely that they will be able to design future generations well.)
pornel 23 hours ago [-]
Heck yes!
I don't get why USA wants to keep losing money on sustaining failed uncompetitive zombie companies. The companies that decided to lose long term competence for short-term gains need to go bankrupt (you've had EV-1, but decided to drill, baby, drill). The greedy shareholders that rewarded destructive value extraction need to lose money, instead of getting a soft exit at taxpayers' expense.
If you want to give a subsidy, give it to something that will modernize and expand manufacturing, not to prolong death of companies whose entire R&D strategy is inventing new subscriptions for old car components.
mrandish 1 days ago [-]
> cheap chinese batteries, EVs, and drones?
The banning or effective banning through tariffs of products like EVs is a pretty dumb economic strategy that rarely works out in the long-run.
verdverm 1 days ago [-]
> Aren't consumers benefiting from cheap chinese batteries, EVs, and drones?
Depends on what country you live in I suppose, but likely a spectrum of yes than any outright no. For example, Chinese EVs are using a different battery chemistry and not putting demand pressure on the more expensive chemistry western manufacturers use
trollbridge 1 days ago [-]
I was surprised to learn Ethiopia is nearly all EV now.
They are a petroleum products importer, so it’s a big win for them.
paxys 1 days ago [-]
It’s the same as the patent argument. If everyone could freely copy everything then yes in the short term prices would drop and consumers would benefit, but over the long term it would discourage investment into new technology because a return would be impossible.
rstuart4133 17 hours ago [-]
> It’s the same as the patent argument.
The Chinese are open-sourcing these models - meaning they are literally giving them away. That's starkly different behaviour from that which drove patents and copyright.
There are all sorts of other differences too. If you develop a drug or publish a novel (which are respectively areas where patents and copyright have strong justification), you have to kiss a lot of toads before you get a prince. Once you have a potential prince drug, it takes years and billions to get it into the market. But right now, AI labs are seemingly churning out new prince models almost weekly, on hardware that will be obsoleted in a few years by hardware that makes finding princes faster and cheaper. In fact they rent the hardware, as Moonshot apparently did here.
When innovation is happening at that rate, patents and other IP restrictions just slow things down. If AI patents are aggressively granted and enforced in the USA, I suspect the outcome would be the same as batteries. China swept the market with LFP, and one reason was because while the USA developed them, there was no competition forcing the manufacturing price down because in the USA the patents weren't re-licensed cheaply. China had the foresight to secure a deal allowing them to develop and sell LFP domestically royalty-free. Natural competition within the domestic market took care of the rest. It looks like they are using a similar strategy for AI.
To me it looks like they have come up with a better formula for using capitalism to drive innovation than the one the USA is using.
But nobody really pays the 3x (ie. api) rates, except for enterprises. Everyone else are using the consumption plans, which are heavily discounted[1], possibly cheaper than even the chinese models, which don't do consumption plan discounts. Even in your linked reddit thread, the OP agreed with this sentiment.
Psh. Enterprises. There can't be more than a couple of them out there.
charcircuit 1 days ago [-]
I have paid API rates. I needed AI to cleanup file space and I didn't want to gamble with the alignment of Chinese models.
applfanboysbgon 1 days ago [-]
For now. We've seen this pattern play out literally a hundred times in tech and you're incredibly naive if you think this will last forever. And what is your point? It should be okay for US consumers if the US government illegalizes accessing open-weight models within their borders because they're currently getting a subsidized token rate?
teravor 1 days ago [-]
the distillation everyone talks about in respect to LLM's isn't nearly as easy as most think.
none of the frontier labs provide probability distributions over the tokens which is the actual method of distillation you use to train a smaller model based on a larger one. they don't even provide all the tokens.
therefore this so-called distillation the frontier labs whine about is just a set of clever methods to work the existing LLM into the training process for a new model. methods like having the existing model grade the output of the new model and work those grades into the RL method. give the new models structured tasks and use the existing model as a source of truth for those tasks and a myriad of other hacks.
efficiency scales with the gap between the models and generally allows an efficient bootstrap process. the implication that distillation wouldn't allow further advancement is false however, you can then start doing the same thing the frontier labs have been doing: dumping cash on humans to provide the signals or burning tokens on exploratory paths and grading the results.
what openai and anthropic don't like is that fact that all the cash they burned can be used to benefit everyone and not just them. and that no matter how much more cash they burn to build up the gap it will closed at a small fraction of the price.
skeledrew 1 days ago [-]
Super interesting. So Fable was really made available... a couple weeks ago? And K3 a few days ago? That's a really impressive feat to distill enough data AND train AND review to get a release that works really well in that time period. Mad props to the Moonshot team :flame:.
tristanj 22 hours ago [-]
Fable was launched on June 9 (for 72 hours), then K3 launched on July 16.
Timeline wise, Moonshot had over a month to post-train K3 on Fable distilled data, which is more than enough time.
skeledrew 17 hours ago [-]
So in 3 days they were able to distill enough data to make a significant difference in the K3's performance? Without triggering any limiters when there's already suspicion of distillation? That's still a heck of a feat. Like the bank robbers who were able to keep coming back to take more even though the bank was aware they had gotten robbed recently and had the resources to put sophisticated security systems in place.
HeavyStorm 1 days ago [-]
Poor AI labs... All they hard earned training, done via scraping a lot of people works for free, now being scraped through payed subscriptions...
MiguelVieira 1 days ago [-]
Here's a site that asks the same questions to 22 models and compares how similar their responses are.
According to these results GLM 5.2 is very similar to Google Gemini and Kimi K3 is very similar to Fable 5.
The American frontier labs are not similar to each other.
xyzsparetimexyz 1 days ago [-]
Interesting. This dooes lend credence to the distillation idea. Good for them!
trollbridge 1 days ago [-]
GLM 5.2 is light years ahead of anything called “Gemini”.
qeternity 1 days ago [-]
I think it’s fairly obvious the Chinese labs are doing mass distillation.
I also think the Fable accusation is wrong and it was most likely Opus 4.8 which itself is likely a distillation of Fable.
sailingparrot 23 hours ago [-]
You have it backwards IMHO, obviously can't prove it, but I would bet that Opus was used to bootstrap Fable.
It becomes very confusing since we have started calling everything distillation, but most likely what both Anthropic did for Fable and Moonshot did for K3 was using Opus traces in the reasoning SFT stage during mid training.
thih9 1 days ago [-]
One of the replies:
> @MehdiKarech
> I don't remember letting Anthropic or Open Ai scrapping my GitHub, my research gate and all my online writings L O L
and the Chinese distilled western cars for years and years before that and weren't shy about it.
HarHarVeryFunny 1 days ago [-]
Of course - buying a competitors product and tearing it down to analyze it is business as usual.
The stupidest part of this is that Anthropic don't even provide the real reasoning traces in their model output. It would be like Ford buying a Chinese EV to tear down, then realizing that the seller had removed the battery and charging system before shipping it to them.
Does anyone believe for a second that Anthropic isn't sending requests to all the Chinese models and analyzing the crap out of them to assess how capable they are, what their reasoning looks like etc?! I guess they'd call that "using" the model, since that sounds nicer.
verdverm 1 days ago [-]
Sandy Monroe has a business product he sells to Big Auto where he tears down cars, creates a bill of material in incredible detail, and expert notes about the process of how the car is made.
If that's not distillation in the auto industry, I don't know what would be. They all seem fine with, and benefit from, this as an industry.
wmf 1 days ago [-]
I think Xiaomi already distilled the Porsche Taycan.
mrhottakes 1 days ago [-]
wearing a lab coat and safety glasses and mixing cool looking liquids in a beaker Oh, I certainly would.
caycep 1 days ago [-]
distillation of wheat, barley and malt is delicious, though!
js8 1 days ago [-]
People tried to distill a lot of things.. even oil.
solumunus 1 days ago [-]
That tickled me!
program_whiz 23 hours ago [-]
To borrow the argument from the apologists:
But everyone learns by example! How is this any different from a person just reading the outputs of Fable, learning, then producing output. Surely reading outputs, gaining knowledge, then producing work isn't illegal, or all art/writing would be illegal.
Funny how that argument seems so vacuous in this situation, yet others find it compelling when justifying the mass theft of art and writing for model creation. In this case the model is "just learning priors" before it "creates its output which is novel", nothing problematic.
storus 1 days ago [-]
I doubt they did any distillation as Hinton defined it (requiring logit access). They most likely ran a bunch of prompts/conversations and captured the results. Those conversations already missed thinking tokens, replaced by some confusing quasi-summaries. Then they took those and ran basic SFT or maybe DPO if they had competing responses. As there is no copyright on the output of AI, I am not sure where is the "covert industrial distillation" part of the problem.
hashstring 12 hours ago [-]
We also have evidence that Anthropic distilled all human info they could get their hands on for the development of all their models.
Distillation should be fair game given the (current) game of LLM training. Yes, as a model creator you probably want to protect against it, but it does make you a hypocrite.
> The developer OpenAI has said it would be impossible to create tools like its groundbreaking chatbot ChatGPT without access to copyrighted material, as pressure grows on artificial intelligence firms over the content used to train their products.
kevincox 10 hours ago [-]
If anything this distillation is more ethical than the original as it is on machine-generated content rather than copyrighted human-labour products.
pandinus 1 days ago [-]
As with many others among these threads I don't see how the timing works out for K3 to have trained on distilled Fable usage. There should be at least a tacit academic acknowledgment of Kimi's own design efforts.
Distillation itself, however, is still clearly valuable - else competitors wouldn't pay so much to their rival on distillation campaigns or try to circumvent anti-distillation defenses.
As for the morality of it, if you paid for the tokens they're yours. It is already understood that you own the output. Seems to me like a variation of ordinary business arbitrage. Providers might object to certain use-cases or intention and try to craft terms around that, but that's hard to enforce at scale.
SwellJoe 1 days ago [-]
"It is already understood that you own the output."
I don't think it's settled that anybody owns the output. There seems to be some question whether LLM output can be copyrighted (and there should be).
I'd rather it weren't possible, actually. I think it's better for humanity if we acknowledge that what was legitimately ingested into these models is our collective commons (and what was illegitimately ingested into these models also shouldn't exclusively profit the people who illegitimately did so). I don't know how that squares with the AI industry recovering its trillion dollars in investment, but I reckon they should have thought of that before.
syrrim 1 days ago [-]
There's no question about it. Llm output is not copyrightable. Which is moot anyways, because training on copyrighted data is completely legal.
throwa356262 1 days ago [-]
In the meantime, reddit is making fun of Opus for "distilling" Qwen:
I would support distilling even if scraping wasn’t legal, I don’t think there is much of a relationship between the two
wnmurphy 1 days ago [-]
It's funny to me that these models were created by effectively "distilling" all available content including the proprietary works of many other people, but now it's a problem that someone is doing the same to them.
You're using available information (copyrighted works, or the output of another model) to train a model to encode the information in a new form. Why is the former not theft, but the latter is theft?
jerrythegerbil 1 days ago [-]
“However, large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable.”
What’s actually happening behind the scenes is that certain inference providers will classify a prompt and it’s re-routed transparently to Anthropic and that’s used for distillation training, only distilling the complicated traces they need, originating from real user prompts and traces. These inference providers are explicitly blocked in the claude cli if you reverse engineer it.
The real picture is that these Chinese labs have figured out how to get exactly what they need, at a high quality, directly from distinct and unique real user prompts.
It’s only “covert” because Anthropic doesn’t like it, while simultaneously being perfectly fine to do.
throw10920 8 hours ago [-]
> What’s actually happening behind the scenes is that certain inference providers will classify a prompt and it’s re-routed transparently to Anthropic and that’s used for distillation training
Uh, no. There are Chinese networks of tens thousands of fake identities specifically to get access to Anthropic models directly.
Don't make up stuff and/or lie to suit a political agenda. It's extremely dishonest.
mrandish 1 days ago [-]
How was K3 trained on data distilled from Fable when Fable was only publicly available in the last two weeks before K3 was released? The timing just doesn't work.
grim_io 1 days ago [-]
So, if it's that easy and fast to "copy" Fable, is it really worth that much in the first place?
Sounds like the opposite of the conversation Anthropic would want to have.
nchmy 1 days ago [-]
We have information that Claude distilled billions of copyrighted, and otherwise-created-by-others, materials for the development of their entire business.
yeodev 1 days ago [-]
The US gov and AI providers when they steal billions of user data, content and media to train their models on: :)
The US gov and AI providers when funny chinese people steal their data to train their models: >:(
clowns
edit: TIL you can't use emojis on HN
jmward01 1 days ago [-]
If 'distillation' means training on outputs then what is the legal concept of ownership of outputs? And, more broadly, is this something that could be skirted by doing it in different countries that have different legal structures? Basically, are they saying they own those outputs, not the companies that paid for the tokens, and only they can train on them? I suspect a lot of companies are saving their token histories and using them to fine tune internal models.
nradov 1 days ago [-]
The legal concept is that LLM vendors can put pretty much whatever they want in their terms of service, and cut off or sue clients who violate those terms. They have the right to refuse service to anyone for any reason (or no reason at all).
zetazzed 24 hours ago [-]
Who will invest in generating data for the frontier of AI if their output will immediately be used to train a competing model? Forget China vs. US, this applies within-country too. After exhausting all the publicly-accessible data on the internet, the frontier labs started spending hundreds of millions of dollars to generate data across a variety of fields. Just look at Mercor doing $1.2bln/year with 90% coming from the top labs (https://www.theinformation.com/articles/mercors-fast-growth-...). That pushes ahead what AI can do in medicine, science, coding, and math. But if other companies are going to free ride on this investment, it doesn't make sense to continue. So AI will largely hit a wall, frozen at the current level and work will all shift to cheaper inference.
sailingparrot 23 hours ago [-]
> So AI will largely hit a wall, frozen at the current level and work will all shift to cheaper inference.
Looks like an ideal outcome to me until (if) we are able to solve the alignment issue.
Alifatisk 23 hours ago [-]
Honestly, I don’t have any sympathy at all. Anthropic can complain all they want, but they seemed fine with pirating books. What Moonshot AI has done is to offer almost Fable 5 comparable performance at lower prices than Anthropic insane margins. This is what I call competition, which the Director seem to embrace.
This is what the Chinese always been good at. Take expensive innovation and streamline it to lower prices. But we are at a point where labs like Moonshot actually contributes a lot to the research field as well. They are pushing the innovation forward and squeezing the prices. Very well done.
Whats even weirder is the bizarre mechanisms Anthropic implemented to prevent distills which they had to sacrifice their customers for. They hid the internal CoT reasoning and returns summarizations instead. This made it difficult for users to trace things. They made Fable 5 silently switched over to Opus 4.8 if it detected blacklisted prompts (almost anything triggered this) to sabotage distills. And now, they are still complaining about distills? So their customers have gotten sacrificed over nothing.
Whats even weirder is the timeframe here, no way the Moonshot team managed to plan conduct a large scale distill, then pre-train, RL, fine-tune, benchmark, marketing and release to their platform since Fable 5 got whitelisted.
> they developed a sophisticated internal platform to conduct large scale distillation
I am very curious about this and would love to learn more on how they did this. Wish we had more details. I know the team behind DeepSeek have also done clever things to distill too. I am aware of these ”transfer stations” that acts as a proxy, but I don’t think they are helpful in this case.
nessex 17 hours ago [-]
Paying the list price Anthropic decided upon for outputs from an LLM, then using those outputs for your work? Or sourcing those outputs from others that paid for the service, and using those? If that's "distillation", it seems fine to me. And a lot like what most are doing with these same providers.
Anthropic: training AI is "transformative", it's not copyright infringement if we don't re-transmit the copyrighted books we trained on
Anthropic: training AI models on our outputs is stealing our secret sauce, outputs that could only be produced by us
Anthropic: AI model outputs are unreliable and do not reflect Anthropic's views, we are not liable if they harm you
(not exact quotes, they're "distilled")
So Anthropic "distills" knowledge of others with reckless abandon, packages it up, sells it to you, claims it's your fault if anything bad happens but then also lobbies to treat you as a criminal if the outputs you paid for end up being transformed into any sort of competition for them. By you, or others that use the outputs you paid for.
cmiles8 1 days ago [-]
But wasn’t fable distilled from knowledge taken from others? I get why Anthropic is angry here, but it would appear they’re not really in a position to complain about this.
softwaredoug 1 days ago [-]
OpenAI and Anthropic should enter into distillation agreements with other US labs. Turn a threat into a profit center.
Other US labs cannot directly distill from OpenAI/Anthropic as it’s a violation of the terms of service. It holds other US labs back. Leading them to build second tier models And in the end OpenAI/Anthropic may be unable to prevent distillation.
Why fight it when there’s clear money to make here?
gozucito 23 hours ago [-]
OpenAI/Anthropic are already charging for access to their closed models. They're getting paid.
They are also not interested in agreements. They want to keep as big a moat as possible because they love money. And you need two to tango.
softwaredoug 23 hours ago [-]
Yeah they may not be interested. But pretending you have that moat a dumb strategy that's not working.
NetOpWibby 1 days ago [-]
GOOD
I love using Claude but Fable's unusable wrt useful work like cryptography, biology, &c.
Kneecapping my productivity when I pay $100/month is annoying af.
zmmmmm 23 hours ago [-]
> We have information ...
Hard to think of a weaker way to express this. Strongly suggests veracity of said information is poor.
MillionY 10 hours ago [-]
As far as I know, this is the easiest way to make that accusation true.
Good. If Fable is really so smart, it wouldn't let itself be distilled.
kamranjon 1 days ago [-]
So here is an important question I think.
If LLM outputs aren't copywriteable and you create your own synthetic training set using Fable and share it publicly on huggingface, and someone else uses that training set to fine-tune a model, would this be considered illegal?
I ask because this happens all the time, synthetic datasets have basically become a key aspect of training a model at this point. I even generated a synthetic set from DeepSeek v4 to aid in fine-tuning a classifier just a few weeks ago.
So I just wonder on what grounds any of this makes sense, I wouldn't be surprised if some of these American labs were using open models on their own self hosted infrastructure to generate training data, but by nature of them being open nobody has to know.
I'll make a prediction: I don't think we will ever see any of the evidence of this "distillation" before they end up implementing some type of ban.
4chandaily 1 days ago [-]
Seems to me like Moonshot is a paying customer, and if their business isn't worth the money Anthropic is charging, perhaps they should raise the price per token charged for it. Otherwise, I don't see why this is a story. "AI Company pays another AI Company for training data" just isn't that interesting.
alastairr 1 days ago [-]
Presumably the frontier labs themselves can do their own distillation far better than the chinese labs can. Why can't they just beat them at their own game and release / host low cost intelligence and own the whole game. There will always be a market for the more expensive frontier intelligence.
econ 1 days ago [-]
I have an idea! If they are so hungry for citable content they should start a cheap or free blogging platform with images and video and a blogroll and verified credentials and resumes, with your own html css etc and domain name and a git server and a mail client and their own advertisement platform and aggregator and a chat platform, scientific journals too obviously, tools for writing and publishing books and documents. API available everywhere to avoid training on its own output.
Because there is no way in hell I'm going to make an effort creating quality content for existing platforms. The website should be entirely my own without moderation subject only to my local legal system.
Can just insert this comment as a prompt and vibe code everything in a few days⸮
Now, of course it's in the "creators" interest to prevent their biz models to breakdown due to piracy but it will be interesting to see how it will turn out. Its similar to any other form of piracy in internet age: you can't pretend to have global distribution and absolute global control at the same time.
blaufast 1 days ago [-]
The frontier labs' work is more akin to discovery than artistic expression. An art piece is valued for its uniqueness and individuality, but AI is valued for verifiable correctness. Discovery cannot be unseen and is easily replicable. I think the AI labs are in a tough situation because their work is more similar to fundamental scientific discovery than say, a unique painting or song.
Mendel doesn't get a cut every time somebody uses the principles of heritability he discovered, and Einstein's family aren't getting royalties if you compute relative speeds. I think the frontier labs should expect to be treated more like scientists than artists in this regard.
neals 1 days ago [-]
How does one distill? Just send a million request asking for information? Start with the letter A?
verdverm 1 days ago [-]
Probably the agent workflow traces, including thinking sections, are of main interest. Used in late training for decision making and problem solving strategies.
NichoPaolucci 1 days ago [-]
I wonder if this points at a “shared” future (or at least things will eventually converge there whether companies like it or not). Ultimately, if you’re going to release these models that are fundamentally built on shared data - it’s pretty wishful to assume you’ll be able to harbor that model and the data, forever, and profit from it.
It also leads me to think about things like the original release of Fable 5, people were complaining that it was safeguarded too much - if you lock the models down too much they cease to be useful. So it’s going to be increasingly difficult to protect a model from competition while ALSO keeping it useful.
nradov 1 days ago [-]
We might see a future where the US frontier LLM vendors place really strict licenses on them. No consumer access. Only sell to enterprise customers in a limited set of countries, with heavy monitoring and auditing down to the individual employee user account level. (I'm not saying that this is a good thing, just that some LLM vendors might try that approach to maintain their "moat".)
alightsoul 1 days ago [-]
how is it possible to distill fable only a month after its release? maybe they are confusing opus with fable.
orbital-decay 1 days ago [-]
Distillation is a superficial step and doesn't need a lot of data, it's not "stealing the model" like they want everyone to believe. 99% of work is already done by that point. That said, it's pretty clear K3 has Claude's data in the training set (either Opus or Fable), as it repeats Anthropic's prompt injections. (not that it matters to anyone besides Anthropic themselves)
wongarsu 1 days ago [-]
If by "distill" they mean "used it for fine-tuning" then they might have used it in the final stages of fine-tuning of Kimi K3. I image they might have already been using Opus, and when Fable became available it was easy to switch over to it
It would have been a tiny part of the overall training, given the timeline
cmdocidjcije 1 days ago [-]
Create a couple thousand Claude max accounts and split the work amongst them perhaps.
skeledrew 1 days ago [-]
That's crazy income for Anthropic to invest into Legendos.
The line "how could they do it in such a short time" is absolutely idiotic. The infrastructure was already there, it's massively parallel, and it's not like the only thing being distilled on is Fable.
sosodev 1 days ago [-]
A month seems plenty long enough. They're not rebuilding the entire model from scratch. It's just getting Fable to act as a teacher model for some of the final reinforcement learning on the base that Kimi already had.
epolanski 1 days ago [-]
I think that would also be a bad idea, as all models opus 4.6 got increasingly smarter, but also crappier at following instructions or genuinely assisting.
They just try to figure out what the goal is and hyper focus on solving it.
Heh, even just telling fable don't commit doesn't work half the times, let alone more complex instructions.
supriyo-biswas 1 days ago [-]
Honestly, it wouldn't surprise me if they just found evidence of distillation once in 2025 against some Chinese AI lab, and they've been lying about the rest to create a narrative.
warkdarrior 1 days ago [-]
I also heard that K3 stole the 2020 election, among other things.
The Irony. These models have been created distilling Internet without ever asking for permission or paying anyone. Internet was the first model.
12 hours ago [-]
tanh 1 days ago [-]
For code can't they distill from public GitHub commits? If they could figure out who used Mythos/Fable assitance in the commits.
SwellJoe 1 days ago [-]
I think they're distilling "reasoning", not merely code. There's plenty of human generated code. What they're trying to extract is the process by which really large models "think" their way through complicated problems. That's what all the "traces" datasets on HuggingFace are about.
speedylight 15 hours ago [-]
It’s ok when they steal the entire internet and everything not the internet they can get their hands on but distillation is bad.
dizhn 16 hours ago [-]
More and more I am starting to get a racist vibe from these claims like they could not possibly have trained a model like the white man can.
muldvarp 1 days ago [-]
Okay? We have information that Anthropic sucked up all of the internet for the development of Fable.
jstummbillig 24 hours ago [-]
I think at some point the issues around copyrighted work and model distillation have to be disconnected to advance either idea.
1) Compensation of right holders is one issue.
2) Distilling models is an entirely separate issue, because model building is value add, and that is important because if we arrive at a place where you can produce a model, that gets to ~100% of what people perceive of the models value (on top of also not compensating right holders, yourself) you are discouraging development of better models and, again, in no way helping with issue 1)
Unless anyone actually distills a model and then also does something for rights holders, any schadenfreude simply detracts from this issue, in addition to the other issue (well, that might not be an issue if we would rather slow down model development right now, but again, forever worse models still don't help solve issue 1)
1 days ago [-]
eyevz 1 days ago [-]
K3 frequently refers to itself as Claude in reasoning when instructed to play a role.
1 days ago [-]
1 days ago [-]
wseqyrku 1 days ago [-]
Every model is distilled internet. This is just the natural progression of that.
Chance-Device 1 days ago [-]
Hmm. I wonder when this was detected. And was the CoT trace cut from Fable from the start on June 9th or just after the export ban and relaunch? Is this what the export ban was actually about? I honestly don’t know, just wondering aloud.
jchw 1 days ago [-]
Is this person trustworthy? I struggle to believe that in the relatively short time Fable was available it has already been distilled so effectively. If this really is actually true, very impressive work.
riknos314 1 days ago [-]
If it's true that in under 15 days of access significant improvements were realized in K3, then the moat of closed-weight models is far smaller than previously thought.
Doesn't bode well for the valuations of these labs.
asadotzler 1 days ago [-]
A company that distills LLMs should be called "Moonshine" not "Moonshot"
Ba-dum-tss
bradrn 24 hours ago [-]
Hah, I had the same thought on seeing the title!
Catloafdev 1 days ago [-]
I wonder how they detect this kind of thing. Seems like this is going to be a perpetual issue until it stops being worth doing.
Side note, didn't they stop releasing real thinking tokens for Fable? Or is it still part of some subs or API usage?
kouteiheika 1 days ago [-]
Assuming they did then they surely paid for them, which makes it "not stealing". Am I also "stealing proprietary U.S. technology" by harvesting my Claude chats from my `.claude` directory and training a bunch of models on them?
That said, I doubt the "they distilled Fable" is the reason why K3 is as good as it is, considering the timelines involved, and that Anthropic hides thinking traces, and their overly aggressive "safety" filters.
This constant FUD spread by Anthropic is so tiring.
sosodev 1 days ago [-]
Model distillation can't be stealing at all if you rationally apply copyright law to it. Anthropic is not deprived of Fable so there is no theft. At best it would be infringement, but even that might not hold up in the courts given the current position that model outputs can't be subject to copyright.
1 days ago [-]
yencabulator 23 hours ago [-]
It's at most a Terms of Service violation escalated into political theater because checks notes China bad.
skeledrew 1 days ago [-]
> harvesting my Claude chats from my `.claude` directory
Just reminded me to set a backup on that directory. Just in case someone sees it fit to override my setting to preserve my chats for 10k years.
goldenarm 1 days ago [-]
I have information that Anthropic distilled the internet for Fable
Stevvo 22 hours ago [-]
Taking statements from this administration at face value is foolish. They have put out so many falsehoods that its safer to assume all of it is false.
sscaryterry 1 days ago [-]
How much credible, provable evidence? None really.
seydor 1 days ago [-]
the bar for credibility in the US administration is in the subatomic scale.
proof is generally not even needed
mrbonner 1 days ago [-]
Hah tales as old as time. what’s next? Distillation of Disney theme park?
1 days ago [-]
woggy 16 hours ago [-]
Great engineering feat considering the small time window that Fable was available for.
qmmmur 19 hours ago [-]
Doesn’t feel nice, does it, Anthropic?
chasd00 1 days ago [-]
if there are no consequences then who cares? You're not going to take Chinese companies to court and stealing IP is nothing new either. It's going to take some sort of policy change at the federal government level to do anything but they haven't done much up to this point. Maybe AI is important enough to actually get some kind of policy change, sucks for everyone else who have had their IP stolen with no consequences whatsoever.
gensym 1 days ago [-]
I fear they are laying the groundwork to ban US citizens from using Chinese models, so that they can make sure that Brockman gets his money's worth.
chasd00 1 days ago [-]
why would you fear that? At least it's something to fight unfair business practices. As for American AI companies violating copyright there's a venue for that, the courts. File a case if you feel your copyright has been violated and get your day in court.
1 days ago [-]
mbix77 1 days ago [-]
Didn't they just pay a fine for stealing all those books?
mrandish 1 days ago [-]
$1.5 billion fine for downloading 7 million books from LibGen and other pirate torrents.
That's also the case where the judge ruled that training AI models on books could qualify as fair use, but storing millions of pirated works in a central internal library without licensing constituted copyright infringement. It will be interesting to see if courts consider training on data distilled from a model fair use. Assuming the allegation is true. Someone distilling data from a cloud-hosted model:
- Paid the model creator to use a publicly available product.
- Never copied or even had access to the model source code or weights.
- Created a derivative work based on the model's responses to their particular input.
- Trained their own model on the distilled output
That distilled output is arguably a collaborative creation because a distiller's prompts are their own unique intellectual property. So they never pirated anything. I'm struggling to see how distillation is copyright infringement. At most it seems to be a paying customer violating one of the license terms, perhaps akin to a "no commercial use of derivative works" clause. But in the case of giving away an open weight model, is it even 'commercial use'?
I guess if the distiller asserts copyright on the weights but gives them away, it's technically 'commercial' but even if they can win that argument, they're left with zero direct damages and suing for some value delta based on the alleged revenue they were deprived of. Is that delta the difference between the distilled model existing and the next best non-distilled open weight model existing? And then they have to collect damages from a portion of the revenue of third parties who commercially served that free model?
wincy 1 days ago [-]
Well I mean it still worked out for them because they wouldn’t have had the 1.5 billion to license before doing the training and the company exploding into a trillion dollar company?
seydor 1 days ago [-]
Good pretext for bannign chinese APIs in the USA and their vassal states.
If distillation is so good, why aren't US companies distilling each other?
exabrial 1 days ago [-]
Oh the irony... LLMs go and read a billion pirated book, but now are crying when their models get distilled from a million queries.
scronkfinkle 1 days ago [-]
so they distilled one of the best models in the world AND released it for free to everyone. Where can I send them flowers as a thank you?
blks 1 days ago [-]
Considering that LLM are created using stolen content or against licensing, it’s only fair to distill and open source them.
nmeofthestate 1 days ago [-]
Weird 'discussion'. Almost entirely single messages with no threads, all with the same anti-Anthropic/AI position.
stratos123 15 hours ago [-]
Yeah, some AI-related posts on HN are strange. It might be real people coming to gloat whenever they see a title they like, but I sure hope somebody is checking whether they are real people.
xinayder 24 hours ago [-]
I wonder if this was caught with the malware code Anthropic included that detects if you're in China...
InsideOutSanta 1 days ago [-]
We already know K3 is really good; you don't need to glaze it even more by telling us it's like Fable.
Biologist123 21 hours ago [-]
Are they serious when so much of training was on illegally acquired data?
Live by the sword, die by the sword.
yanhangyhy 12 hours ago [-]
even they are using Distillation, its also a pretty scary ablity.. with such short time and result.
guybedo 1 days ago [-]
i have information Anthropic distilled thousands of books, articles, etc ... with their author consent.
paradox242 22 hours ago [-]
They stole all the data for the model in the first place so fair is fair.
K0balt 19 hours ago [-]
lol?
All your base belong to us?
Crocodile tears?
I’m less upset by this than I am about “music piracy”. And to be clear, I’m not upset about music piracy.
codedokode 1 days ago [-]
Do you by chance also have information about Anthropic's training data sources?
chriswunan 1 days ago [-]
Human only have this much knowledge. They will become similar anyway.
stephbook 1 days ago [-]
I don't know what purpose these "they copied us" crying is ever going to achieve. Europeans stole Chinese silk worms. US stole European books, looms and rocket scientists. Who cares? Be grateful you've got people inventing stuff worth copying.
8note 19 hours ago [-]
if these can be distilled so easily, does the model actually need to be so big and trained with so much electricity?
cloudie78 14 hours ago [-]
Okay, and?
We have information that all the big labs used copyrighted works for training.
The big boy labs wanna cry now about distillation? Training an LLM is distillation too.
Or are they crying because they don’t actually have a moat and they want Uncle Sam to step in somehow, lest the entire bubble pops and economy unwinds?
Here’s a lesson from the automotive industry, people want econoboxes not formula 1 cars.
Grimblewald 22 hours ago [-]
And I have evidence that anthropic distills from openai, moonshot, alibaba etc. So unless anthropic holds itself to the standard they're implying should be followed, why should I care others did to them as they do unto others? Seems like a nothing burger. Has the same vibe as a bully crying foul because they got hit back. Also, if distillation is what got K3 to where it is, why is it better in many areas? Also, big fucking kudos to k3 team for putting this together do fast givne how new fable access is, if it is true and had a meaningful impact. The real story there for me becomes one of extreme competence.
kjs3 1 days ago [-]
People who built a business model around stealing other peoples stuff are vewy, vewy upset that someone has built a business model around stealing their stuff.
Oh no! Anyway.....
nozzlegear 1 days ago [-]
Cry about it IMO. Anthropic reaps what they sow.
1 days ago [-]
tacone 1 days ago [-]
So it is as "dangerous" as Fable?
laweijfmvo 1 days ago [-]
if its that simple, why doesn’t Anthropic just distill its own models and release Fable 5.1, 5.2, ...?
stratos123 15 hours ago [-]
Distillation lets you train a comparable model for a fraction of the cost (it's effectively a way to get a lot of very-high-quality training data). If you're already at the frontier, pushing it requires the ordinary, expensive kind of training.
guess_who_is 1 days ago [-]
If you ask fable, it will identify as deepseek
sajithdilshan 1 days ago [-]
How the tables have turned. It's okay for Anthropic to train their models on copyrighted data, but it's wrong to steal the stolen data from Anthropic models.
m_ke 1 days ago [-]
Anthropic should think hard about all their fear mongering. It will only end up backfiring on them and everyone else involved.
They definitely used closed private saas products to train their own models, to prove that just drop random small screenshots of any popular product behind a login screen and see how well it's able to identify all of them. ex: https://x.com/michalwols/status/2079968211865330165
"Well, Steve, I think there's more than one way of looking at it. I think it's more like we both had this rich neighbor named Xerox and I broke into his house to steal the TV set and found out that you had already stolen it."
phs318u 23 hours ago [-]
Do I, as a paying user, own the output the LLM has produced or am I merely licensing the output?
Most paying users assume ownership, in which case I’ll do with that output what I want.
If the LLM outputs are not owned by the user, but are actually licensed, please clarify the terms of commercial use.
nahuel0x 1 days ago [-]
Information wants to be free.
benterix 13 hours ago [-]
LOL. A thief accusing someone else of theft.
Also, I love their choice of words. Like "distillation against", "stealing proprietary technology" - it's all aimed at certain people.
bparsons 1 days ago [-]
IP protections for me, not for thee.
dudeinhawaii 23 hours ago [-]
I think most of the comments in this thread are missing the point. It's not about whether it's legal/ethical/etc.
It's about the narrative that "Chinese models are at Fable level". The truth (if correct) is the China continues to copy, and the proprietary US Models continue to lead the state of the art.
There is no K4 without Fable 6, GPT-6. That, matters.
sailingparrot 23 hours ago [-]
> There is no K4 without Fable 6, GPT-6. That, matters.
That's simply not true though. Chinese labs very clearly have the entire stack developed and working. Using traces from claude allows them to shorten their training time by some amount, that's it.
Remove Fable 6 and you still have K4 eventually, just 2 months later at best.
dudeinhawaii 22 hours ago [-]
[flagged]
deaton 1 days ago [-]
Who cares. Anthropic distilled the entire internet, and then a good bit more beyond that.
dmitrygr 1 days ago [-]
OMG someone used our data to make an MK model! Just like we did to every author in the world!
rambojohnson 1 days ago [-]
who cares. all these frontier models are trained on theft.
martinjc 1 days ago [-]
So what? I want the best model at the cheapest price. You guys illegally trained on books, movies, audiobook etc.. Why should we care?
cregy 1 days ago [-]
I think the major take away is, Chinese labs are very good at stealing others AI work and offering it did a fraction of the cost.
This will be the reason the AI stock market bubble bursts.
Unless they add in protection similar to patents, which stops these copied models being used by business.
BigTTYGothGF 1 days ago [-]
Good for Moonshot.
mattrighetti 1 days ago [-]
Is distillation something we have to live with or are there ways to prevent it?
sosodev 1 days ago [-]
Realistically you can't prevent distillation. OpenAI / Anthropic are slowly moving towards hiding the steps in-between input and output (hidden thinking), but that only helps so much. Imagine you put a file into Claude and say "do X to this" and it returns it to you without showing any of its internal reasoning. That's harder to distill, but the simple mapping of input to output still creates very valuable training data. It is reflective of all the training the model did to learn how to do that transformation.
make3 1 days ago [-]
You can also get it to think in the output tokens pretty easily, eg "Here's a math problem, I want your reasoning first, then the answer" which is what I assume they're doing.
Cytobit 1 days ago [-]
You make it sound like a bad thing.
HarHarVeryFunny 1 days ago [-]
You can prevent it by outputting a reasoning "summary" instead of the actual reasoning trace.
Which Anthropic already do.
sowbug 1 days ago [-]
If you build a device that can help build devices, you shouldn't be too surprised when people use it to build devices.
recursive 24 hours ago [-]
To the extent that we have to live with AI and LLM, I believe we have to live with distilling also.
epolanski 1 days ago [-]
If it was genuinely useful, we would've long reached the point where you train a model on a previous one's output in an ever improving loop.
But this doesn't actually work.
4ndrewl 1 days ago [-]
So? Their business model requires building on the labor of others for free. Isn't that how it works?
xnoto 1 days ago [-]
"we ripped off the entire ecosystem of copyrighted data but I draw the line when we get ripped off"
NDlurker 1 days ago [-]
Good. Keep it up
caycep 1 days ago [-]
honestly if they did what he said they did, it seems like it would be cheaper just to train your own model from the get go
browningstreet 1 days ago [-]
I haven't seen that point yet, and I was looking for it. Presumably Moonshot paid for that Fable access and Anthropic got paid. How much of the frontier model revenue stream is supported by paid distillation traffic? Obv paid kimi services are eating that on the other side, but money is changing hands at every stage.
Gud 1 days ago [-]
OK, DIRECTOR Michael Kratsios, but why should we give a shit?
American AI corporations are pushing up the prices for computing, making it unaffordable for the common man. Additionally, they have built their entire business on stealing(yes, stealing) work from us.
So fuck em
stranded22 1 days ago [-]
Seems like an advert for K3 to me.
Fable level performance, for much lower price.
But really, this is the USA getting ready to bring AI companies completely under the control of the Trump administration for ‘national security’
sensanaty 1 days ago [-]
For a buncha supposed capitalists, they sure do hate fair competition eh?
juancn 1 days ago [-]
So?
We have information that Fable was distilled from humans.
If it works it works. Isn't that the argument?
AI outputs are not copyrightable, so distillation is fair use.
It may be a TOS violation, but that's a private matter. Cancel the accounts used for distillation and be done.
matheusmoreira 1 days ago [-]
So what? Am I supposed to be upset by this? Distill away. Actually, can I help out somehow? As long as they keep publishing open weights, I'll give them my full support.
ProofHouse 1 days ago [-]
didn't Anthropic distil a few tens of millions of books?
zuzululu 1 days ago [-]
If distilling is fair then so is banning it. It's funny how people cry about Anthropic using books and materials to train itself but then when Anthropic does something about it they think its unfair. Pick a lane.
stldev 1 days ago [-]
So company who stole stuff to make their stuff is mad because another company is stealing their stuff.
And now a regime best known for lying to their own people is the one trying to convince me?
Go, China!
bakugo 1 days ago [-]
Fun fact about K3's distillation:
As of a couple months ago, when using Claude to write adult content through the API, sometimes it will silently inject a system prompt giving the model a bunch of guidelines on exactly what kind of adult content it's allowed to write, steering it away from anything "questionable" ("Claude will not write etc etc").
Moonshot distilled Claude so hard recently, they actually ended up distilling this prompt injection, too. Using K3 to write adult content results in it randomly hallucinating the injected Claude prompt during thinking, and it will quote parts of that prompt, complete with the name "Claude".
Not that I think distillation is a bad thing, just thought this was funny.
noncoml 1 days ago [-]
Yes, I know this is not Reddit but Clarkson’s “Oh no! Anyway…” is the perfect, and most fitting, reaction to this. Nothing else to say
xingped 21 hours ago [-]
God I'm sick and tired of AI companies being so damn whiney all the time. Shut up already. No one cares.
solumunus 1 days ago [-]
Get your violins out folks.
tamimio 1 days ago [-]
“If you can’t compete with them, get them banned”
- US AI companies
make3 1 days ago [-]
There's no real way to compete with someone who gets the output of your own work for almost free in comparison.
I sympathize with the argument saying that they ripped the whole Internet and books first though
superloika 1 days ago [-]
I think they deserve, by Justice, to have their models pillaged and raped, just like they did to the internet. They didn't ask for permission when they took the entire of the internet, after all, and given their behaviour is nefarious, it's of Justice that they receive nefarious treatment by others, including chinese AI labs.
The Chinese are not gonna deterred, but the posturing by the Americans is so blatantly hypocritical that everybody is cheering for their demise. See, for example, one of Francis Fukuyama's latests videos on youtube.
cwmoore 1 days ago [-]
I still believe taxing the bots, and implementing actual UBI, would address both problems.
jgilias 1 days ago [-]
Will the UBI apply to everyone worldwide? As that’s where the original dataset came from.
cwmoore 16 hours ago [-]
Will it?
runarberg 1 days ago [-]
And I believe a socialist revolution in international solidarity of the working classes against our exploiters the capitalist owning class would address both problems as well (and more), but in the meantime I’ll be happy whenever I spot poetic justice in the wild.
Chance-Device 1 days ago [-]
I think we have some historical precedents for this that went less well than you might hope.
runarberg 1 days ago [-]
I am not aware of any revolutions where an international solidarity of the working classes overthrew their capitalist exploiters. I am aware of a couple of national solidarity movements which overthrew a monarchy (France; Russia) and a dictatorship (Cuba), but no international ones.
Chance-Device 1 days ago [-]
I think it’s safe to extrapolate.
1 days ago [-]
surgical_fire 1 days ago [-]
First: Even if true, I don't care.
Second: Post is rich with allegations but light with evidence. Can very well be bullshit.
cvanelteren 21 hours ago [-]
I mean it identifies as Claude in 1 in 5 cases when asked so pretty obvious to me.
xiaodai 15 hours ago [-]
imagine if a company's MO is to PAY Claude to use their model and distill the results. That's fair use of their model!
1 days ago [-]
traceroute66 1 days ago [-]
"we have information" says a US Government official who almost certainly has had Anthropic and/or OpenAI on the phone spinning him stories.
See also, don't trust anyone in Trump's government who says "we have information".
"they distilled us" is fast becoming standard US FUD.
The same as people telling me with a serious face that the Chinese models are distilled just because it says "I am Claude".
I am not the only one, look at this post on interconnects about Kimi K3 for example:[1]
It should be clear looking at this model that if adversarial distillation from the closed frontier models in the U.S. contributed, it is at most to a relatively small degree. AI observers who followed the distillation panic and came away with the wrong conclusion that Chinese AI labs are only producing good models due to IP theft are in for an awakening – that Chinese companies are extremely good at building models in the same way the leading American companies are.
The problem is… what are you going to do about it?
This is obviously an idiotic and dangerous Cold War and has no happy ending.
1 days ago [-]
cyanydeez 1 days ago [-]
oh know, better make them pay up your legal fines for stealing all that from the public good.
mikelitoris 24 hours ago [-]
“Hey stop stealing my data. I stole it fair and square.”
overgard 1 days ago [-]
It's both acceptable and inevitable. Play stupid games, win stupid prizes. This is what they deserve for the awful way they treat creators and creative people's intellectual property and livelihood.
lowbloodsugar 1 days ago [-]
China has done this with absolutely everything, starting with “customs inspections” of ships engineering sections by “inspectors” drawing diagrams of what they see. Bit late to be worrying about it now. This only matters now because China is now near parity in tech and vastly superior in production ability. Meanwhile we run out of bullets in a five month war with Iran.
voxelc4L 18 hours ago [-]
Watching the tenants of intellectual gatekeeping bang up against the inevitable paradoxes of late stage capitalism - a spectacle worthy of jiffy pop.
felooboolooomba 1 days ago [-]
The pot calling the kettle black.
GiorgioG 1 days ago [-]
I don't even understand this fucking AI "arms race" at all. Who can burn money the fastest?
drop_star 1 days ago [-]
America, the perpetual victim
mrhottakes 1 days ago [-]
We're winning so much, we're getting tired of winning
pbgcp2026 13 hours ago [-]
... but these MFs cripple our work by introducing stupid "guardrails".
CamperBob2 24 hours ago [-]
"No fair stealing what we stole fair and square"
avazhi 1 days ago [-]
And?
Nobody cares. This is neither a controversy nor news, and that would be the case even if Anthropic hadn’t just settled a 1.5 billion dollar lawsuit where they trained Claude on thousands of books without permission lol.
To be clear I’m not taking a jab at OP - I’m saying the labs crying about distillation have neither a legal nor a moral leg to stand on. There’s nothing wrong with distillation.
jauntywundrkind 1 days ago [-]
Two recent ones that really really hit me,
> we're entering the most geopolitically volatile moment since the trinity test lit up the alamogordo desert and the only US policy prescription is a big button labeled sinophobia
> every vendor cranking the big dial labeled "sinophobia" and looking back at the us government for approval
The government itself doing the propaganda here, skipping the vendors. Sinophobia intensifies. War drums of "be afraid be afraid be afraid" beat louder.
It's so bad, it's so stupid. Kimi lands one showing pretty clearly this was absolutely the determining concern happening at vast scale, that they can just a lot of this themselves, and this noise pollution from the most hopelessly lost aggro administration ever still gets blared out the trumpets of war & discord. What a joke. Give me a break, give it a rest.
War here is less winnable than the Iran war they started. They're going to make America itself so much worse, these people so hungry to put down free and good models. This pathetic attempt is not going to work, you are just going to once again hold the US citizens hostage & make their lives worse, for sick political games.
orangecat 1 days ago [-]
Not surprising that Bluesky is perpetually in peak woke mode, but the racism claims are absurd. If Russia were doing the same thing, would we have no problem with that because they're white?
jauntywundrkind 1 days ago [-]
This feels radically off target to me.
Sinophobia as such (not always, and there's obviously a relation) isn't about the Chinese people, about their racial identity. It's about a menacing aggressive world power (actually two such MAWPs in this particular example), about a tired anti-Communist McCarthyism (which has always been an excuse to clamp down on the left/progressives, a menace to free speech).
If Russia were still the USSR and the cold war hadn't ended and they were our AI competition, & where giving away the latent matricies of reality that the US profiteers extracted by stealing all the worlds knowledge illegally, and which they want to use to raise the ladder & leave a permanent plebeian underclass, we the US imperial fatcats would be doing the same sabre rattling and fearmongering and tension raising for sure. Sinophobia here is just a particular fear of "other" for the only other that's relevant.
And that sucks, no matter who it is we are trying to other here, no matter what phobia the propagandists of the GOP and Technofascism are spinning, ginning up. To fixate is to be unable to see the point.
tibbydudeza 1 days ago [-]
Proof - they also claimed that China has an ASML UEV machine - crickets when ASML said it was impossible due to all the safeguards and assistance needed to operate one.
The current US administration is known to be collection of BS artists and liars.
TZubiri 22 hours ago [-]
The discussion of
"haha LLM companies stole data and now they have their data stolen so it's the same thing and it's fair." was reductionist when it started, and it's been like 3 months, and every internet user throws it like it's the hottest take ever, have another take please.
Also have nuance, don't jump to hit your HOT_TAKE key in your keyboard, actually read what the chinese are doing, and then you can pass on your judgment on whether it's ok or not.
It's not the same thing if they scrape an openly published dataset and it's an IP dispute. Or if they are using,network and financial pooling mechanisms that are shared with CSAM providers and cybercriminals, mutually providing each other alibies, and using black markets of passport-backed identities to setup thousands of accounts and circumvent bans and detection.
While we are at it, if there's a case that was settled, it's a closed case, it can never invalidate any other disputes. That case is closed, and it was settled by the parties that claimed to be damaged, that's done. If you didn't think so, you wouldn't have taken the settlement, and if you didn't have a say in the settlement, it's because you weren't damaged so who cares, go make a claim where you are the defendant if you believe otherwise. But thankfully in no legal system does the existence of a claim against you prevent you from making claims of your own.
Nuance is a good thing.
sleepyguy 1 days ago [-]
Is this a surprise, I think history has proven that the Chinese technology theft is part of their strategy. They let the American tax payer or "The West" shoulder the cost and then steal it.
Waiting for the whataboutism....junk away...
Perenti 22 hours ago [-]
Yes, China stole the recipe for pasta and noodles from America. China stole gunpowder from America. China stole printing from America. China stole the idea of paper money from America. China stole the idea of entry exams from America. China stole the invention of the rocket from America. Seriously, do you really believe this American Bubble bullshit that everything of value was invented in the USA?
buellerbueller 1 days ago [-]
It wasnt until 1891 that America extended copyright protection to foreign authors.
IP "theft" has been a longstanding part of any developing nation's economy.
treetalker 1 days ago [-]
rules for thee but not for me
dang 1 days ago [-]
Plenty of HN readers feel this way and it's a good point, but it has also become an entirely cliché response which pops up like mushrooms anytime "distillation" appears. That means it's against the site guidelines, which ask:
I don't mean to pick on you personally! It's just that reflexive responses always tend to show up first in a thread, when what we really want are reflective responses [1]. Similarly, there's a strong tendency for threads to turn into generic discussions, whereas what we really want are specific ones [2].
To be fair, that's also the case for the link itself we're discussing.
supriyo-biswas 1 days ago [-]
I think then we should ban these sorts of posts about the allegation of distillation, since being able to post the story but then warning accounts with comments about the hypocrisy, is not the correct way to go about it.
dare944 1 days ago [-]
In the controversy over AI and intellectual property rights, the notion of asymmetrical enforcement of those rights is presently an active and dynamic rhetorical position. If you suppress responses espousing that position, how am I to judge the breadth of the position in the debate?
When I see a lot of the responses, even cliche ones, it tells me people are passionate about the issue. Right now I think that's an interesting piece of information.
Bratmon 1 days ago [-]
Laughing at the idea of distillation being bad is exactly as cliche/flamebaity as complaining that your model got distilled.
No more, no less.
cmdocidjcije 1 days ago [-]
At this point distillation is part of the ecosystem and everyone should embrace it. If distillation is a threat to one’s business model, then the business strategy needs to shift.
phikappa 1 days ago [-]
sure but anthropic is not literally in the comments complaining, so it's not quite apples to apples, right?
cassianoleal 1 days ago [-]
I have to agree. I've flagged the post.
latexr 1 days ago [-]
I agree in the abstract, but perhaps the way to avoid generic responses is to disallow (or segment) generic submissions. This website is no longer HN, it should be renamed AIN. There is only so much to say about the subject, and if cliché submissions keep getting accepted and upvoted and shoved to every visitor without a way to avoid them (barring leaving the website entirely), then people will eventually gravitate to the same responses. If your neighbours play loud music every night, they don’t get to complain that everyone is always mentioning the loud music to them.
You are a fantastic moderator, but there’s only so much even you can do. If nothing changes about the website, the problem will only get worse. I warned years ago that this would happen, the signs were on the wall immediately.
Der_Einzige 1 days ago [-]
You just told on yourself big time about being a lapdog for the US Feds. Easily one of the worst moderation decisions you’ve ever made and that’s impressive given your track record.
Telling someone to have more than 20 characters in a comment on HN is bootlicking?
Absolute lunacy.
buellerbueller 1 days ago [-]
Requiring someone to have more than 20 characters, when 20 or fewer will do perfectly? Yes.
unethical_ban 1 days ago [-]
Let me put it in an HN-acceptable format:
Given the disregard for intellectual property rights the AI labs had in creating the technology, many people feel no sympathy for second-order AI labs using similar techniques to build technology off the US frontier labs.
I think fighting distillation will always be cat-and-mouse, and that it's more of a concern for the stockholders and perhaps an iota of national security. It can't be stopped entirely; the "problem" will always be there.
I'm much more concerned about asymmetry of power between citizens and their governments with omnipresent surveillance and analysis being done on everyone living their lives. Societies throughout history have taken as a given their power to overthrow malicious governments when things hit a breaking point, and I am scared that this technology will lock societies into a state of total subordination for eternity.
onraglanroad 1 days ago [-]
That's a better comment but
> Societies throughout history have taken as a given their power to overthrow malicious governments when things hit a breaking point,
simply isn't true. People throughout history have simply accepted that society is the way it is and sometimes used whatever means they could to get to the top.
Revolutions have been very rare and usually ended up with the revolutionaries simply taking the place of the previous rulers. "Meet the new boss..."
unethical_ban 1 days ago [-]
Then let's call it a reset, one way or the other. I'm afraid society won't be able to hit the reset button on a government in the future.
062570864389 1 days ago [-]
[dead]
OhNoNotAgain_99 23 hours ago [-]
[dead]
redsocksfan45 1 days ago [-]
[dead]
jasonmp85 1 days ago [-]
[dead]
beaker52 1 days ago [-]
[flagged]
dang 1 days ago [-]
> /giphy nobody cares SpongeBob meme
Can you please not do this here? There's nothing wrong with it, we're just trying for something else on this site.
"Don't be snarky. [...] Omit internet tropes. [...etc...]
Genuinely want to know if you're surprised that there have been almost no commenters concurring with (what seems to be) a well-substantiated (& on the face of it defensibly center-right) tweet from the WH
Will contrarian dynamic take off here?
Will it be flagged? Will you unflag it?
Edit: catigula, mattrighetti
With more substantive technical comments towards the middle
dang 24 hours ago [-]
Yes, I think it just takes time for more substantive comments to show up.
You may not owe AmericanAIBros better but you owe this community better if you're participating in it.
stego-tech 7 hours ago [-]
I’m familiar with the guidelines and I stand by what I said, including how I said it. This has been a regular talking point from AI companies since DeepSeek first hit the scene, and I feel a glib response to what is very clearly an insincere and nakedly hypocritical talking point is warranted after years of this slop.
Hypocrisy doesn’t warrant professionalism, it warrants corrective action; in text, the best I can offer is a tone and tenor that matches the original argument. Considering these dolts have now made the claim that open weights somehow equates to AI communism, this sort of response is even more necessary than before to reflect the complete absence of decorum from the people making these grievances in the first place.
mbmbn 1 days ago [-]
[flagged]
codedokode 1 days ago [-]
Smartphones are mostly Chinese now (except for iPhones and Samsung).
strictnein 1 days ago [-]
Apple and Samsung are ~40-50% of the global market, depending on the source. Saying that smartphones are mostly from Chinese companies isn't accurate. And Samsung and Apple are gaining market share, while the major Chinese brands are losing market share.
Foxconn facilities in cities like Zhengzhou (known as "iPhone City") and Shenzhen.
strictnein 3 hours ago [-]
Yes, but I was responding to this comment:
"Smartphones are mostly Chinese now (except for iPhones and Samsung)."
Which implies that we are talking about the companies, not where they are manufactured.
goldylochness 1 days ago [-]
[flagged]
strictnein 1 days ago [-]
[flagged]
sailingparrot 1 days ago [-]
Yes they distill, but if you think you can trivially get a frontier-level model by "just" distilling from Claude's public API. you fundamentally do not understand the amount of work that goes into a modern post-training stack.
Without even talking about the fact that any distillation that was done was on Opus, as the timeline of Mythos/Fable vs Kimi 3 release dates just do not match in any plausible way.
“Distillation” is just a term of art these days. Usage of competitors’ models these days is mostly around RLHF (minus the H, I guess).
strictnein 1 days ago [-]
Very interesting, thanks!
anfogoat 1 days ago [-]
> Do we need 40 people saying the same thing about how they don't feel bad and it serves them right and all that surface level stuff on every single one of these?
Apparently we do, given that we've got govt officials wasting time, money, and effort on whining about this now.
jrflo 23 hours ago [-]
Agreed, it's becoming a bit of an echo chamber anytime Chinese models are discussed here. This is really par for the course for China though, in physical manufacturing they've been doing this kind of thing for decades, I'm not surprised the mentality has persisted into the AI race. Looks like rather than jumping ahead, they will remain persistently 6 months behind.
recursive 24 hours ago [-]
The responses will continue until distillation-posting improves.
If we keep hearing accusations about how someone distilled something from someone, it seems like one of the few reasonable responses.
If my friend keeps complaining about how the inside of his car is wet, I will probably keep telling him to close the windows when it's raining, even if he thinks that is a tiresome take, lacking in insight.
calendar938 1 days ago [-]
Let's be real. In reality, it's the Chinese kid getting the 1600.
strictnein 1 days ago [-]
lol, true true. Should have used a different example.
oybng 24 hours ago [-]
Who cares. It's all theft
alexruf 1 days ago [-]
Who cares? Is it theft if a thief gets robbed of their stolen goods?
Gives me more of a modern Robin Hood vibe tbh.
No, seriously: first of all, that's not the AI labs' data, it's ours. And if the AI labs think they can rake in tons of money using our data, then I'm actually glad if someone comes along and at least offers us a good product at reasonable prices.
brap 1 days ago [-]
Anyone surprised by this is incredibly naive.
By all means use whatever works for you, I’m not even going to try to make an argument on ethics (and honestly I’m not even sure where I stand, given the behavior of American AI companies).
But I just cringe every time I see people acting like any of this is done in good faith.
Open source coming out of China is a state-sponsored criminal enterprise, built only for the benefit of the Chinese regime, one of the worst to exist in human history.
There are millions of samples available on huggingface and models explicitely trained on output produced by fable. There has been no action taken against them.
Another example is that it appears that the upper limit of what you can do is ultimately dependent on people working on the model, otherwise grok would be a LOT more competitive pre-cursor acquisition.
And lastly, kimi architecture is vastly different than that of fable as it uses mechanisms developed by... kimi themselves. US AI labs are inspired by opensource advancements just as much as open source labs are inspired by traces from models such as fable.
Claiming in any shape or form that fable disillation is one of the primary reasons why kimi k3 is so competitive is slandering the work of other labs that cooperatively push the open-source models forward.
edit: (moved this to bottom) The only argument they have here is that they use GB300 GPU's which for some reason should not be available to chinese citizens.
Serious question.
I'm from the US, and I think it's hilarious.
"Thank you," the KGB says. "We do our best but truly, it's nothing compared to American propaganda. Your people believe everything your state media tells them."
The CIA agent drops his drink in shock and disgust. "Thank you friend, but you must be confused... There's no propaganda in America."
In a world where information sources are only going to dwindle, it is not in anyone's interest to empower actors that will use these to manipulate perceptions
Can you name one? It's an honest question, first of I'm not American, but also the stuff I do come up with (asking an LLM how to blow up a school or whatever) would also be handled similarly in China, so those would be a wash, and I can't think of any that aren't.
Another thing I would try if I had access to the models and enough proxies to hide behind is asking for advice on software/movie piracy or seeing to what extent the models can be elicited to straight up argue against the validity of intellectual property, though there it seems more probable to me that the US models would be permissive.
- The Trail of Tears
- The Tuskegee syphilis study
- Use of Agent Orange in the Vietnam War
- Open Air biological warfare testts in civilians eg. in 1950 San Francisco
- The only use of nuclear weapons against civilians?
- Coca Cola and "american culture"
- Neoliberalist economy
- Spreading blame for their sins to other "white" nations
plus one: The text input method to HN comments :(
you are comparing a wooden stick to a fighter yet, try again
Of course that might change in the future but as long as the Chinese companies continue publishing their research and models it only makes it easier for third parties to catch up with them.
To avoid future situations where money is invested on hype only with no regards to what societal disruption it causes.
Raw materials vs. Value add.
They are different things, like ore and metal.
Distillation is a new thing we need to understand, it's probably closer to IP than not.
We're talking about things like text people wrote, not some kind of raw data floating out in the ether.
Ore has value, a different kind of value than the output of the refinery.
https://arxiv.org/abs/1503.02531
although i doubt there has been a legal case over it yet in the context of the legality of stealing shit but IANAL.
It's completey insane that we still don't know how Open Source would work, that the laws are vague and we're still technically waiting for the courts to decide on cases.
The government should a) legislate and b) create test cases and run them through the courts so that we can have clarity.
you mean like the regulatory clarity surrounding stealing shit to make the LLMs in the first place?
> [There is] extensive litigation on the limits of Fair Use to AI development. Currently, we only have 3 first instance decisions out of the 53 cases being tried. It will likely take a decade before we understand how Fair Use applies to any one step in AI training, let alone all.
https://www.britishcopyright.org/wp-content/uploads/BCC-Fair...
> The government should a) legislate and b) create test cases and run them through the courts so that we can have clarity.
if so, it would be nice if they approached the instances of stealing shit chronologically. but that's just my view.
It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ...
... but Chinese SOTA foundries directly using distillation as fair game.
I don't think there is any coherence to any of these arguments - other than 'we liking big companies'. That's the only common thread.
What is more reasonable:
- There's some grounds for fair use by SOTA models to ingest content, so long as they are not reproducing it ... very roughly speaking.
- SOTA makers are producing novel works, there is value add in that process, again roughly speaking.
- Distillation is a bit of a grey zone, producing random content as arbitrary input is one thing, but producing training sets is another. I think there's a coherent line in there somewhere, I'm not sure where it is.
Anthropic's models simply do not give you their reasoning output - they give a sanitized "summary" instead, for this exact reason, so that the output is not useful to anyone who might want to use it for training.
You can't distill what you are not given - simple as that.
Are Chinese using the output of US models to help create some additional training data for their own in some way? Yes - quite possibly (e.g. LLM as judge), but its got nothing to do with distillation.
I think where the definition may be be invalid, is in the creation of 'unrelated data sets for training' models, for unrelated issues.
Creating training sets that mach a models core training, is definitely distillation, it does not have to expose the reasoning traces.
Synthesizing data for some arbitrary thing ... I'm not sure that would be the same thing.
It's hard to draw the line.
But the Chinese models are absolutely distilling - and would not be competitive without this distillation.
At the same time, there's a lot of real innovation and regular building going on at the same time over there.
I don't know why it's so important to you to use the word "distillation", but it's the wrong word to use.
BTW OpenAI on twitter also said that Kimi 3 "cannot be explained away by distillation or anything like that". The timeline of how long it takes to train a model and when Fable was released don't even line up. This is just Anthropic as usual trying to manipulate the US government into helping them shut down competition.
This isn't really a debate, I'm not making a fine point - just check with all of the various defintions of the term.
Moreover - the 'reasoning traces' are not required for distillation at all.
Finally - it's entirely possible for them to have used Fable for later stage fine tuning.
It's fair to be skeptical of Anthropic (and everyone else) - but this is 'distilling'.
Are reasoning traces required for distillation? Well they are if what you are trying to distill is reasoning, such as coding expertise.
Do you need reasoning traces for "LLM as judge"? No, but it would be highly perverse to call that distillation when there is a more accurate name for it - LLM as judge.
If you want to call use of Anthropic's redacted model outputs in any fashion that violates their terms of service (using them them to help develop anything that competes with Anthropic) as "distillation" then I can't stop you, but it reduces their claims to a joke.
Finally, as noted, OpenAI (who are just as anti-Chinese as Anthropic) said that Kimi 3 can't be explained via distillation (even true distillation!!), or even "anything like it". But random internet guy, you, disagrees. OK.
This is just straight-up factually false.
The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, making it even more valuable.
There's absolutely nothing about the distillation process that requires that reasoning in the first place, either. That's a definition that you made up.
Chinese models are, factually, distilled from Anthropic models. I've personally repeatedly asked several different Chinese LLMs what their name is, and they answered "Claude".
Don't make stuff up to suit a political agenda. It's extremely dishonest.
I'm curious what you are doing to get them to override their own name that they were trained on and/or have as part of their system prompt?
I'd assume that the Chinese are scraping the internet for training data the same way western companies do, so for sure there will be a lot of AI generated content in their training data - you don't need to be paranoid and assume they must be getting it all direct from Anthropic.
Useful for what is the question. Nobody is debating whether the outputs of LLMs are valuable.
Given that Anthropic have redacted their true reasoning, and replaced it with a "summary", specifically designed to be useless for distillation purposes, it would certainly be highly ironic if this summary was in fact "even more valuable" for that purpose as you are claiming!
Useful for distillation. Any employee at a frontier AI lab will tell you this. This is known in the industry, and it's an open secret that some US labs (OpenAI) distill on the others. Again - don't just make up stuff for a political agenda.
> specifically designed to be useless for distillation purposes
No, it's designed to give feedback to the user, in a way that minimizes its value for distilling. It's still valuable, and so there's a good chance that they'll remove it entirely as a result.
> it would certainly be highly ironic if this summary was in fact "even more valuable" for that purpose as you are claiming!
I did not claim that. Read my comment again:
> The output of a reasoning model is immensely valuable even without the sanitized summary of the reasoning process - that Anthropic's service does expose to you, making it even more valuable.
Because apparently I have to spell it out:
The output of a reasoning model is valuable, even if it didn't have the reasoning summary. Anthropic's models have a reasoning summary. The reasoning summary makes the output more valuable than if it didn't have a reasoning summary. It does not make it more valuable than having the full reasoning.
Here's the thing: no-doubt a summary, if it at least reflects some of the logic connecting response to request, is better than nothing, so this can still be useful additional training data, but a model trained on it would be learning to generate these summaries, not the original withheld reasoning, so "distillation" seems an intentionally emotionally-wrought way of describing it (the "summary" is generated by a different smaller model - not the one the rest of the response came from).
It does bring up an interesting point though - RL training in general results in "cargo-cult" reasoning - you train a model to follow the steps (mistakes and all - Karpathy) that got to a result, without understanding why they worked. If this training on summaries works just as well as training on detailed reasoning, then it just highlights how having a few breadcrumbs to follow/regurgitate is all that it takes.
At the end of the day, without having internal logits or original reasoning traces, "distillation" (which suggests one model being derived from another) just seems a very manipulative way of describing this. OK, so it's a terms of service violation - a customer is using Anthropic model outputs to help create something that competes with Anthropic, but that's it. They are not copying Anthropic - they are, one assumes, using other models to generate cheap training data that they would otherwise have to pay people to generate.
This is why people are calling out the hypocrisy - Anthropic are apple-pie American innovators when they appropriate other people's copyright data for training, but Kimi are evil communists when (we assume) they use data generated by Anthropic (not even copyright protected) to help train their own.
To be fair I find it hard to take this too seriously, shouldn’t it be trivial to just replace “Claude” with any other string in your “distillation” dataset?
There are open-source Deepseek and Qwen models - "distilling" doesn't involve breaking terms of service or hitting an API because you can literally run local inference or even just inspect the weights directly, and that's intended because they're open source.
It's categorically different for a nation-state to build massive illicit networks of fraudulent identities to do distillation over tens of thousands of accounts to intentionally bypass providers' terms of service, intention for their models, and business model that very explicitly proprietary and not open source.
https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...
If Claude did distill on proprietary PRC LLMs - then fine, shame on them - I condemn that and I expect others to do the same. But there are no open-source Claude models. The only way for PRC models to have those responses is if they distilled Anthropic's models from their APIs.
> To be fair I find it hard to take this too seriously, shouldn’t it be trivial to just replace “Claude” with any other string in your “distillation” dataset?
...and what would happen when it read all of the books and articles about Anthropic and replaced "replaced Claude Opus" with "replaced Qwen Opus"? Did you give any thought to this at all before saying it?
The fraud part and using stolen accounts or credit cards or blatantly violating the terms and conditions (i.e. reselling subscriptions not using outputs in certain ways somebody might not like) is indeed categorically different.
Using uncopyrightable outputs of an AI model obtained legitimately to train your model is not inherently interlinked with any of those things. I don’t really see how the model being proprietary or “open” is particularly relevant when talking about the outputs.
Even using the word “distilling” in this case is deceptive and biased. It implies that the Chinese are somehow stealing Anthropic’s models or their weights and somehow directly transforming them into new models. That’s certainly not what’s happening in any direct sense.
e.g. what if I agreed to send all my Claude code session files to Deepseek or whoever? There would be nothing wrong about that since I and not Anthropic own those files and can do whatever I want with them. Using certain different ways to obtain them of course could be highly illegal.
there's no creative work between the weights and the tokens being made.
whats the big deal if chinese companies sell an exact replica of the model? its a summary of a variety of works of text and images
If not there is not there is no grey zone whatsoever.
for the record, i've always been rabidly pro-copyright since i worked at a performing royalty organization (prs for music) circa 15 years ago, way before i joined hn.
i don't use llms for that reason.
> It's also sadly hypocritical to see all this rhetoric here on this thread decrying SOTAs for using a variety of content for their inputs as somehow sTeaLINg sTufF! ...
when i see a spade, i call it a spade. just because the US has utterly stupid copyright provisions that are wide open for abuse, i.e. fair use, doesn't mean abusing those provisions at scale is morally acceptable.
> ... but Chinese SOTA foundries directly using distillation as fair game.
two wrongs don't make a right, but the irony is at least something.
> I don't think there is any coherence to any of these arguments - other than 'we liking big companies'. That's the only common thread.
the corpos can get fucked as far as i'm concerned.
> What is more reasonable: ... There's some grounds for fair use by SOTA models to ingest content, so long as they are not reproducing it ... very roughly speaking.
*only in the US.
There is no second “wrong” here.
Model outputs are not copyrightable. I think that was already established?
Or do you think that Anthropic should own all the code generated by Claude? Surely that would be somewhat problematic?
If Anthropic feels that some of their customers are breaking their EULA (nothing to do with copyright infringement though) they are free to stop doing business with them. Maybe even sue them in civil court for breach of contract (again nothing to do with copyright infringement though)
anthropic are essentially saying in this tweet they believe a moral wrong has been committed against them -- "unacceptable behaviour" etc.
plenty of people have been vocal about the fact anthropic have committed moral wrongs at scale in building the products in the first place, with the question of legal wrongs still being worked out.
so, two moral wrongs. legally, fuck knows.
I mean you are right in a way of course, it’s just a matter of degree and perspective, though. If one thing is moderately morally wrong and the other is potentially lightly morally wrong I don’t think it’s fair to equate them.
To me the situation is a bit like Google coming out and saying that its morally wrong for someone to build a competing open operating system on top of Android while stripping all Google services and “stealing” their ad revenue. Just seems silly and hypocritical.
I'm just nothing that HN rhetoric is contradictory.
But this:
"when i see a spade, i call it a spade." -> this is anti intellectual absolutism.
If it were some true injustice, then fine, but that is clearly not the case.
There is ample room to contemplate that even copyrighted works could be considers fair use as training material.
"the corpos can get fucked as far as i'm concerned."
Ok that's fine - but then don't expect anyone to respect your principles if you don't have any other than 'screw that group!'.
I'm sympathetic to it (!!!) - but if we want to call a 'spade a spade' in a legitimate way, then we can do it in consistent and principled way.
Periodic reminder that HN is not a collective or a singular entity and is actually a bunch of different people with different opinions. Often the people with the loudest opinions get upvoted to the top - and often the "side" represented at the top is different from thread to thread.
And the result is force feeding an AI slop generator with a subscription while making personal hardware 3x+ times more expensive.
No wonder people are fed up with this behavior.
You do realize the 'investors' are the one's who 'own' companies and therefore the IP?
One person can a corporation be.
> ... but Chinese SOTA foundries directly using distillation as fair game.
As someone who says it’s fair game, it’s less that I’m being hypocritical and more that I don’t care that one thief had their shit stolen by a second thief. I also wouldn’t care if someone distills the Chinese models. It’s just thieves all around and if they want legal protection or moral outrage from the common man then my view is that they should stop stealing first.
LLM outputs are not copyrightable (or rather the user is effectively the only one who can own it). It would be problematic if Anthropic owned all the software generated using Claude..
If what Anthropic/OpenAi did for training is theft then the Chinese models also are a form of theft. If they didn’t steal then I don’t think the Chinese firms did either.
That entirely settles it and there isn’t much else to say about.
If Anthropic feels that other countries are violating their EULA well they are free to stop doing business with them.
They claim LLMs "uncopyright" their inputs. So if I take, say, 50 Mickey Mouse comic books, tell ChatGPT to read them and produce 50 "Buster Beagle" comic books that there is ZERO "copyright contamination" and I own those 50 output comics without Disney having any claims on them whatsoever.
Or if I ask ChatGPT to "make a spreadsheet software like Excel, Sheets, Calc, ..." that, again, there is zero copyright claim possible from these people.
It has not been tested, of course.
That would stun me, but it's a little hard to read.
ed: to clarify, I totally agree that a huge chunk of the value in LLMs is coming from the source material. My point was just that training an LLM takes more resources and expertise than distilling from an existing LLM so I don't think the equivalence between training and distilling is entirely justified.
It's the most CS-major take ever!
Whether we think they're paying enough is another question, but "I'm paying for content so can protect it" doesn't seem inconsistent.
We may decide that giving models away for free means they don't have to license content (judging by HN comments), but currently that doesn't seem to be the case as Meta is facing lawsuits for its open models.
(Obligatory stratechery piece: https://stratechery.com/2026/whos-afraid-of-chinese-models/ )
The same principle can be applied to distillation - it is a fair use. You just shouldn't use illegal ways to access the models being distilled.
To the commenter below: if it is illegal - has the police/FBI report been made? Otherwise it is just a civil court matter.
It does seem to be becoming the norm for AI companies to licence premium content in America, judging by the deals they're making. It doesn't seem to be done by the international distillers. It's a cost that American open models will seem to have to pay but not international.
This is a very surprising claim to me (and I imagine many small website owners who keep getting scraped by Anthropic and OpenAI).
Do you have a source?
https://digiday.com/media/a-timeline-of-the-major-deals-betw...
I don't really understand why you think they're relevant, given that this conversation is about the training itself.
Even the news orgs say explicitly in the press releases that it's about training on their archive
eg. http://ap.org/media-center/press-releases/2023/ap-open-ai-ag...
---
edit, examples:
Wiley https://newsroom.wiley.com/press-releases/press-release-deta...
Shutterstock https://investor.shutterstock.com/news-releases/news-release...
Axel Springer https://openai.com/index/axel-springer-partnership
Stack Overflow: https://stackoverflow.co/partnerships
Disney (for characters in video. Video is especially where licensing is a big difference internationally right now) https://openai.com/index/disney-sora-agreement
etc.
The news corp one had a leaked price ($250mill), so they don't seem to be insignificant. These would have to be included in API prices I presume.
International distillers doesn't use that premium content, so they don't pay for it. They do pay for their access to the models they are distilling. Thus providing the revenue stream to those models. Thus those models make profit off the content they used for training. The content they mostly have't paid for.
>It's a cost that American open models will seem to have to pay but not international.
It goes both ways - American companies and their business are protected by American laws and have access to the market protected by those laws, etc.
This doesn't seem to be true. They are training on their own scraped data overwhelmingly (we can extract copyright data from, eg, deepseek). They couldn't get nearly enough tokens through the American APIs to train a model on alone.
> American companies and their business are protected by American laws and have access to the market protected by those laws
Absolutely. Currently international providers are selling inference on the American market though, I don't know how that will sit legally the way things are currently going.
In any case the laws are being written now, but I doubt these will have worse protection than software does, which has far better protections than copyright
Software is protected by copyright. Some software may also be protected by patents, but last time I checked, AI generated output of any kind was not patentable.
Also note that the OpenAI/Anthropic argument is that the model training is sufficiently transformative to satisfy the fair use of the original content for training.
By that same argument, when distilling the distillers aren't using the original content the OpenAI/Anthropic models were trained on - the distillers are interacting only with the "sufficiently transformed" content of the OpenAI/Anthropic models and are normally paying for that.
There is also that old phonebook rule that facts can't be copyrighted. So, if i asked the model about bunch of phone numbers, i can publish the resulting list, can train my model on it, etc. Such approach doesn't allow to reproduce copyrighted works of course - and as we know the AI output isn't copyrightable, so it looks like basically any output i get i can use whatever way i like.
Also, if model output distillation is shown as some form of reverse engineering I assume the DMCA can apply
I agree that you can't patent a book, but I would point out that you can patent an idea, which may only appear in a book or journal article.
For example, a patent describing a chemical process. The actual idea of how to do it is public domain, go look up the patent. Print it out. Do whatever with those words. Its fine. Building a plant to go do that chemical process to make that same output chemical in that same way, that's IP infringement. Its not the words, its the idea.
Who are you saying owns that IP? The people who trained the model? The people who ran the model? The people who wrote the prompt? The person who paid for all of that to happen?
If the model output is owned by the person prompting it and paying for the tokens, what's the problem here?
If the model output is owned by the trainer of the model, that's a big nasty can of worms.
I mean otherwise it’s a very slippery slope, effectively it would give Anthropic the ownership of any code generated by its models..
MBAs and non technical managers = inept Catbert-type charlatans.
Software engineers, devs, etc = geniuses capable of mastering any domain, innate ability to be right on any topic.
Or maybe they're going through an intermediary "transfer station" that's breaking terms of service:
https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...
Anthropic's own copyright infringement could apparently be forgiven for 1.5B USD after all, so maybe there's a price that breaking the distillation clause for is acceptable too. Or some other arrangement.
Why is Anthropic's ToS any more binding than that of a rabidly-anti-ai literature blog with 50 readers?
https://en.wikipedia.org/wiki/The_Pile_(dataset)
Another reason is that if you can download a web page without agreeing to a ToS, I'm not sure that counts as one?
I mean i know you know the answer: anthropic is a corporation with lawyers on retainer, and that's really all that matters
> Why is Anthropic's ToS any more binding than that of a rabidly-anti-ai literature blog with 50 readers?
Although I will say, this whole comparison stuff really doesn't seem to be your thing; might impede your analysis quite a lot: https://news.ycombinator.com/item?id=49013148
Maybe ask Claude?
There is no world in which me vacuuming the entirety of human knowledge to make a genai model is ok but hoovering my model answers is not. The hypocrisy is stunning and risible.
Now if you go and make a model based on purely synthetic data and not a single work made by humans, you would have a valid point.
No argument here, I completely agree.
> There is no world in which me vacuuming the entirety of human knowledge to make a genai model is ok but hoovering my model answers is not.
I disagree with this though. Clearly LLMs owe a huge debt to everything that has come before, but surely you'd agree that the models that are produced are something substantial and new and novel which didn't exist before and have lots of value in their own right. Let's be a bit reductive and pretend Moonshot had just outright stolen the weights from Fable somehow, clearly that wouldn't be contributing anything really new or novel. Now of course they've distilled rather than stolen, but the point is similar: how much value have they added along the way?
It's ok for me to use your source code for free as long as I then let others also use my source code for free.
> There is an even higher level of creativity in creating books, songs and all sorts of art used in model training though.
He claimed there was more creativity in model training than in model distillation. That makes no claim about the relationship between the creativity in model creation and art. Why are you continuing to attack a claim that was never made, after a sub thread very explicitly clarifying that that claim was not made?
>Perhaps even more importantly, the current frontier LLM models are self-admittedly the product of enormous quantities of copyright infringement and even less savory inputs, so calling them out for distilling the fruit of that tainted tree reads as highly hypocritical at best.
Context is important. And in this context, their argument only mentions creativity when it belongs to an AI lab. That omission is the blind spot I pointed out. Bottom line is whether or not Anthropic are being hypocritical and yes, they most definitely are, regardless of any attempted sophistry.
There is a reason courts want you to tell "The whole truth" and not just "the truth".
the same argument - a level of creativity in the world knowledge creation that ins't present in the model training on that knowledge.
Or in other words - model creation and training is just a distilling of the world knowledge.
If you think that the addition of a less creative process (model creation) to a more creative corpus ("art") is problematic, then it follows that you should think the addition of a less creative process (distillation) to a more creative corpus (a model) is also problematic.
Like look, I'm not a native speaker, sure. But I think when someone says "value add", that means there was value there (which you claim they're rhetorically erasing), and then that was added to. Under no interpretation of this phrase do I get an erasure of prior value.
So certainly, as long as words mean anything, no, they absolutely did not say or suggest what you claim they did, and what you extract a thus unreasonable amount of obnoxious schadenfreude from, while throwing in a cheap insult for funsies at the end.
It's the second time I feel compelled to reach for this just today: https://i.kym-cdn.com/photos/images/original/002/659/979/108...
The LLM output, is not the same as the input - there is value add.
Of course works used as raw inputs to LLMs required work and are reasonably subject to IP concerns - but they are different.
It's possible that the LLM makers 'owe' the content creators that created the content they used to make their products - it's an interesting but separate question.
We could very well end up where content IP is protected, LLM output is not and visa versa with reasonable legal founding, doubtful but plausible.
But it's debatable if that's the case.
Google stores copyrighted content and produces in in their product.
Also - it's fair game to use snippets of things here and there, if the derived work is novel, which I think it is for LLMs, mostly.
I do agree though, that we ought to draw the line somehow.
How, and why?
> We could very well end up where content IP is protected, LLM output is not and visa versa with reasonable legal founding, doubtful but plausible.
That is the current state of legal rulings - LLM output is public domain, not copyrightable.
Our current laws simply weren’t built for this and I expect the legal status of LLM output is not going to be resolved until Congress actually legislates on this topic.
How, and why?"
How are they even remotely the same?
They're not even used the same way.
One is raw data input, the other is training content - designed to train LLMs.
One is a set of IP derived for other purposes entirely, and has esablished IP law - how you can use someone else's creative work or not ... for LLM outputs, less clear.
Writing books, building Wikipedia, and answering questions on online forums takes a lot of resources and expertise that scraping didn't. So at the very least, we're already one rung down the "maybe you should've asked" ladder.
EDIT: just wanted to add that resource optimization is usually where the contribution of Chinese labs is, so you shouldn't reaad the above parenthesis as a negative comment.
It took me a year to write a book. It took OpenAI and Anthropic a fraction of a second to ingest it. Do you understand now why I give zero shits if it takes Anthropic a billion to train a model, and Moonshot 10k in API cost to distill it?
This is not automatically true. Training and distillation use the same underlying infra and method and there is no intrinsic differences in between.
If that is the whole point you need to clarify why this is the case on an objective level.
I would say building a comparable model using any means necessary (just like what Anthropic and OAI did) at a lower cost is actually more valuable to soceity and Monshoot is arguably generating more value with less.
If the distilled model is cheaper, then it's just LLM's getting LLM'ed.
producing the entire body of human knowledge that Silicon Valley companies absorbed like the Borg did not just take more resources but also a fair amount of blood and sweat, certainly more than the LLM so on that front that comparison also seems entirely justified.
The recent announcement that AI-assisted research produced a counterexample to the Jacobian conjecture--a long-standing open problem in algebraic geometry--shows the original value AI can create. The result was not copied from a textbook; it emerged from AI learning from existing material, much as a human does, and then applying that knowledge in a new way. If that's a violation of copyright, then a human doing the exact same thing would be a copyright violation too. But it isn't.
True, if the human's access to the book was legal
A great deal of training was on the open web, no one should complain.
But at least Meta and Anthropic were caught red handed taking copyrighted works, illegally, for training
I think international IP laws are too strick and onerous, but they were broken to train these models
It’s massive copyright infringement.
The human buys the books.
I wouldn't want to live in a world where technology or general people's wellbeing was held back by obsolete laws that ended up lingering on just to protect undeserving special people at the expense of the rest of society. Remember guilds for tradesmen? They were also a monopoly given by the government to special people. They had their purpose but nowadays we have different ways to keep tradesmen working effectively like license requirements and insurance.
Just to be clear, I think we do still need copyright, but that we might be in a transition period where it has to be redesigned to adapt to AI.
We will all be sorry when professionally written and edited works disappear. An author has a reputation and the incentive to protect that reputation keeps standards high.
> protect our first-party products from abuse like bots, scraping
Won't you think of the trillion dollar corporations?!
>Claiming in any shape or form that fable disillation is one of the primary reasons why kimi k3 is so competitive is slandering the work of other labs that cooperatively push the open-source models forward.
If the distillation is irrelevant to why it is competitive, why do they do it then? Obviously is helps improve their benchmarks/performance to some degree, otherwise they wouldn't need to do it.
Although I will reiterate the fact that distillation is not the primary reason why these models are performing so competitively.
If Kimi k3 really were above Fable 5 then there invariably the USG would have to consider their restrictions on model capabilities excessive, or one would have to admin closed source models are held to more restrictive safety standards than open source models.
>Although I will reiterate the fact that distillation is not the primary reason why these models are performing so competitively.
How would you know this? How could you ascertain exactly how much performance is attributable to their unique engineering/research? If they really were so competitive they could surely make a model that isn't dependent on distilling Fable or other frontier models.
Kimi specifically relies heavily on reasoning traces which is largely due to their training strategy and will perform poorly when thrown into a conversation from another model. Another fun advancement is that they simply ctrl+c ctrl+v'd attention which means that the model can steer where to look in the context window without ever producing an output token increasing token efficiency and attention accuracy as a side effect you end up with weaker prompt adherence.
None of these 'issues' manifest in US models which proves that kimi has diverged and is achieving these capabilities seperately from the architecture that US labs rely on.
I would agree with you during the Deepseek R1 era, but US labs were heavily inspired by open research at that point as well so I wouldn't give them too much credit.
In other words: even if the US (somehow) denies them access to the current OpenAI/Anthropic models, they'll be able to improve based on what they already have.
Actually, that has already happened in many domains, it's just that most western people (USA especially) won't admit it.
at the same time, I don't buy the idea that distillation is unimportant in assessing what Chinese labs are capable of. If it wasn't, why did Kimi's release timing coincide so well with Fable's launch?
and if Anthropic hadn't released Fable, would we have Kimi today? If the answer is no, then I think that's still a very important point to consider.
For some reason folks seem to think that China can take action and then other countries can't also take action or respond to that action and it comes up again and again. China has hypersonic missiles! Pack it up boys time to go home. Nothing we can do. Dang shucks. China distilled American AI models, welp time to just close it all down and let's just write off those trillions of dollars and all the literal geniuses financing and building these things. Oh well China can just copy American models while we spend all the money! Ok we just stop developing models and we'll just copy their models. China will flood the market with their cheap products! Nope can't do anything like, oh, idk, not buy any of those products or just raise the prices on them in local markets. It's never-ending. I don't understand the lack of capacity to reason about other actors that takes commonly takes place. And that's just China, never mind other general issues.
That works both ways, competition and performance spur new developments. You don't think the American labs are looking at Chinese research on how to reduce compute per token?
Let's not forget how much people talked about "prompt engineering" before Deepseek mainstreamed the idea of thinking mode which is now universal
I know this isn't exactly a scientific test, but I had a local Qwen 3.6 27B model implement a fairly sizable feature today. There were a couple of bugs, mostly around me not giving sufficient specifications, but they were ironed out quickly when I pointed it out. I was able to ask the model to create instructions so next time it doesn't fall into the same pitfalls, and it did a great job. 27B local model! (And it was super fast too).
I ran Fable 5 as a code review and it didn't really have any significant corrections.
I guess my point here is that, for most work the frontier models are probably overkill anyway, and improving on overkill in a way that raises prices significantly is probably not a winning strategy.
The only place I can think of where the super high powered models are "required" is if you want to do a ridiculous token burn like GasTown where you just have it run un-monitored on very long tasks. To me though, that's an experiment, not a real workflow. And the way these labs are like "oh we made this (broken) thing in a week using just agents!" always also follows with "and it cost $100,000+ in tokens!". Like, ok, I get it if you're doing research but that's the salary of an entire person.. that can actually learn and improve.
The answer depends on whether you think the AI researchers at Chinese labs are (or can be) as smart, motivated, and as good at math as those working at US labs - a not-insignificant proportion of whom are Chinese nationals.
Note that Chinese companies are free to rent from GB300 clouds internationally. There are large datacenter hubs in Singapore and Malaysia serving chinese and other customers.
Though there is also reported [1] significant smuggling of Nvidia chips into China as well.
[1] https://epoch.ai/publications/chip-smuggling
US dominance is also important for approaches to safety, especially political approaches. If the frontier models are all US-based, safety might be tackled via internal US policy. If other countries can independently train competitive models, international cooperation is required.
Edit: It is also important for the business model. Companies won't be able to justify tremendous training costs if competitors can replicate their product much more cheaply via distillation.
I don’t think China’s necessarily any better, but I’d rather have the most powerful models be open rather than under the exclusive control of the US executive.
https://en.wikipedia.org/wiki/Wolf_warrior_diplomacy
https://en.wikipedia.org/wiki/Chinese_police_overseas_servic...
No, at least not outside of this forum.
We all mostly think these models and US policy are going to drive the exact opposite of that. Wealth will continue to get extracted and funneled to the top, and the rest of us are going to be left with the scraps and left to die while what little social safety nets we had continue to get eroded away alongside losing our jobs.
No, no we do not.
We already know it's false because you would have hundreds of competitors if it was that easy.
The reason why these Chinese labs are releasing good models is simpler, they have access to a tremendous pool of talented people.
Also, companies that use distillation may be competitive but seem unlikely to surpass the companies that are training these models from scratch.
Then why is it a problem?
Another serious question.
Trying to get my head around what the root of the objection is here. There must be some fear, but if that fear is not a fear of being surpassed in the market, then what is the fear?
If I wanted to argue that it's a problem, I'd just say that companies investing billions in training frontier models should reap the rewards. And distillation is essentially theft.
The word I would use is inevitable. It reminds me of the (PC) clones wars…
They claim it because Anthropic are planning to push for protectionism. They just doubled their political spending to $40 million for the midterms to "push for AI regulation" Gee, I wonder what it is they are lobbying for. Certainly won't be OFAC sanctions right? ICTS import controls?
US GOV, under lobbying pressure from Anthropic and OpenAI are going to go full protectionism and restrict Chinese models, I'd almost be willing to bet money on it. They can't really enforce for individuals, but they can definitely tell US based hpyerscalers they can't host them, make it illegal to host the weights, and government procurement restrictions.
I don't follow events closely, but the US has constantly flipflopped on what sort of GPUs the Chinese are allowed to have, not in small part because much of the AI boom's valuation is based on demand for US-made hardware, for which the Chinese have inexhaustible and well-financed demand.
So even this feels a bit hypocritical to me, but my understanding is that Chinese native AI hardware is getting good enough that labs dont feel a huge disadvantage by being forced to buy at home, even if they'd have preferred to buy US chips.
Which is a situation that was manufactured by the constant thread of having their access to advanced GPUs revoked.
I am waiting for a precedent on this one. In general, training on copyrighted material is legal, there is a lot of precedent there. But every now and then there is a case where the owner of the training material wins.
I don't remember the details but I believe one of these instances was when one company trained its AI on the knowledge base of another company and turned it into a competing product. Fair use was denied because of that direct competition. Distilling a LLM to make a competing LLM looks kind of like this, or maybe not, I don't know.
It would make sense for distillation to be legal in every way, LLMs are built on a broad interpretation of fair use, but sometimes, law is weird.
You're making a fundamental assumption: that model outputs are subject to copyright. In the US that's only the case if a human is part of the creative process:
https://www.copyright.gov/newsnet/2025/1060.html
> It concludes that the outputs of generative AI can be protected by copyright only where a human author has determined sufficient expressive elements. This can include situations where a human-authored work is perceptible in an AI output, or a human makes creative arrangements or modifications of the output, but not the mere provision of prompts.
It's showing that 'distillation' is a viable way to reclaim all of what they stole and hoard, and with enough luck their debts will come due in time for them to feel it.
Because what they want them to think is "the AI factory has unique proprietary technology that cannot be replicated"
What they don't want them to think is "it's relatively easy once you know the basics to bootstrap to near SOTA and so the commercial case for selling inference has an extremely short profitability horizon with little if any brand loyalty or lock in".
All LLMs are trained on the corpus of humanity's knowledge, the legacy of everyone who's ever lived and our civilization as a whole.
Anything that prevents or circumvents the accumulation or gatekeeping of this knowledge and puts it in the hands of more people (that are not AI company shareholders) is a good thing. Whether that is done by open sourcing the model weights, the training set, or by making the output better and cheaper, it is all fair game and is, as another poster mentioned, inevitable in the long run.
Correct, but it at least helps answer the question of "how do they make such good models for a fraction of the price???" The answer is someone else spends the untold billions and Chinese labs do a little tweaking.
> Distillation is not illegal by every definition of the word
Note that Anthropic (and USG) alleges [0] not only that Kimi was distilled, but that they actively circumvented measures intended to stop distillation. There are multiple ways that's illegal, including:
- Civil breach of contract. Anthropic's TOS explicitly say you can't do what Kimi is alleged to have done.
- Economic espionage: 18 U.S.C. §1831 criminalizes obtaining a trade secret through theft, fraud, or deception while intending that it will benefit a foreign entity.
- Trade-secret misappropriation: if Anthropic could argue industrial-scale querying reconstructed proprietary aspects of Fable (like by showing it produces similar outputs, as others have done) then it's illegal under 18 U.S.C. §1832.
- California computer-access statute §502 bars knowingly accessing a computer system and, without permission, taking, copying, or using its data.
- Computer Fraud and Abuse Act protects against the case where restrictions against an activity are circumvented (like Kimi is alleged to have done).
> There are millions of samples available on huggingface and models explicitely trained on output produced by fable. There has been no action taken against them.
A lack of prosecution does not make something legal. There is also the scale/commercialization thing, which isn't an issue with random tiny HF datasets/models. Remember: Kimi also sells K3 inference.
> kimi architecture is vastly different than that of fable
How do you know that? Do you work for Anthropic? Also, this has nothing to do with architecture, we are talking about data.
> US AI labs are inspired by opensource advancements just as much as open source labs are inspired by traces from models such as fable.
Cool. The difference is that one of those things is legal (because they chose to open-source) and one of those things is illegal theft of trade secrets (because it was stolen).
> Claiming in any shape or form that fable disillation is one of the primary reasons why kimi k3 is so competitive is slandering the work of other labs that cooperatively push the open-source models forward.
1) this has nothing to do with other labs, just Moonshot (and Z.ai, MiniMax, DS)
2) slandering or not it happens to be completely true, so, there's that
[0] https://www.anthropic.com/news/detecting-and-preventing-dist...
In this way, it is different from literal theft. Stealing money/objects from a thief and keeping them is not justified.
An analogy might be a baker stole 20% of the flour used to bake their special bread, which was then stolen. Both thefts are obviously wrong and bad.
Being a bit tongue in cheek, one could also argue that by releasing their models, Moonshot is contributing much more value than Anthrophic. Did Prometheus not create an immense amount of value, when he took fire from the hands of the Gods and gave it to humans?
"Judge approves a $1.5B Anthropic settlement over pirated books used to train the Claude chatbot"
https://abcnews.com/Technology/wireStory/judge-approves-15b-...
In this case: resolve the theft claims against the US frontier labs, and only then let them make claims against third parties. It would be totally unreasonable for (say) OpenAI to extract a settlement from Moonshot and use that to pay its own claims. Ordering matters.
It seems like you think theft is bad everywhere except when Anthropic does it.
Same thing here. This whole situation is just comical.
Now, if that kid were to print the bootleg translation and sell it to schoolmates, that's worth a slap on the wrist. The kids willing to pay would likely have paid for official copies.
When these LLM labs download our works, feed them into their models, and sell the output to people that used to pay for our work, that's worth a very hard slap. I honestly have less of a problem with the open models.
When I said "Does this matter?" I specially meant that distillation in itself, the data you get from distillation is first and foremost not owned by anthropic nor is it copyrightable. If a user willingly gives up their anthropic reasoning data/traces that is 100% legal no matter what the "terms of service" say as it's not enforceable and would fall apart in court.
And what I explicitely pointed out that focusing so much on distillation is an attack on open research and claiming that the majority of advancements are thanks to US labs which is simply not true (at least not anymore this was somewhat true during deepseek R1 era), but that in itself was inspired by open research.
> How do you know that? Do you work for Anthropic? Also, this has nothing to do with architecture, we are talking about data.
Because anthropic would be the first ones to make that information public and the architecture is unique to kimi... They made it, they wrote papers on it, it's their research.
P.S. none of the quoted laws apply here since no trade information is stolen, the one about circumventing distillation protection might hold up in court although unlikely.
Agree, and this is exactly what Anthropic is alleging.
> data you get from distillation is first and foremost not owned by anthropic nor is it copyrightable. If a user willingly gives up their anthropic reasoning data/traces that is 100% legal no matter what the "terms of service" say as it's not enforceable and would fall apart in court.
It's important to note this is NOT what happened. Anthropic was able to trace data directly back to employees at the company: "We attributed the campaign through request metadata, which matched the public profiles of senior Moonshot staff."
> none of the quoted laws apply here since no trade information is stolen
There is a lot of work showing Kimi models produce similar outputs to Anthropic models, which constitutes trade information. This is not dissimilar to past and ongoing IP suits against Anthropic and OpenAI by showing the models would recreate images of Mickey Mouse/NYT articles etc.
For the record, I'm a researcher myself and I'm well aware how competent the researchers are at the open-source labs/how much they've contributed. But that's not at issue here, my disagreement with you is specific to your arguments about legality; you're conflating what you think should be legal with what actually is legal.
Everything else is simply justifying why it shouldn't, the specifics don't really matter as there is no legal framework to stop china from continuing to distill models and anthropic has proven they cannot use software solutions to stop it either as distillation is still a problem. But I do still believe it wouldn't hold up in court either way as stopping companies from generating training data which was trained on the entire human knowledge corpus is just stealing from thieves and making it 'open' once again so the argument only gets weaker.
edit: to add, the mickey mouse / nyc was because anthropic trained on LICENSED works, not apple to oranges. The original work it was reciting was licensed and not licensed BY anthropic.
This is true, but Kimi also has a variety of defenses. Kimi can't raise unclean hands if Anthropic systematically violated others' terms of use, but it can raise copyright misuse (which is similar in some respects to unclean hands) as well as lack of standing to enforce restrictions in the contract due to the third party beneficiary principle (i.e., Kimi would argue that Anthropic cannot sue Kimi for derived IP that rightfully belongs to third parties whose terms of use were violated by Anthropic, and the proper party to sue Kimi, if any, would be those third parties). That latter argument usually fails in small-scale cases (ProCD) but has been successful in larger ones where the alternative would be anticompetitive.
Ah yes, I remember when Anthropic crawlers abided by the TOS of the websites they slurped up.
All your other points are downstream from this, which makes them pretty tenuous. Labs don't think that ToS or other explicit wishes of content providers apply to them, but they expect everyone else to abide by theirs.
US and CA law really don't care that Anthropic violated IP law elsewhere.
Plainly who gives a flying fuck. The US can claim whatever rules they want and so can China or any other country. On international level all those rules are artificial constructs unless they can be enforced. China can just say for example that they do not recognize copyrights /patents / whatever so it is "legal" for them.
1) its not illegal (it is)
2) it shouldn't be illegal because Anthropic stole training data (thats not how the law works)
I am a practical man. From what I see laws are mostly for common folks and often do not even serve real justice. The higher one goes and the amount of money / power involved the more the laws bend and on international level the only law that matters is the size of one's club and willingness to use it. And when the country with supposedly biggest one starts crying I find it laughable.
The reason why the United States government is weighing in is because it's in the national interest of the US to have supremacy in "AI".
Legality or lack thereof is one of many data points about whether a thing is noteworthy.
Moonshot performing distillation is rational from their point of view. Reducing costs is in the interest of businesses. It's also rational for frontier labs and the US government to add obstacles to this process.
As consumers this is probably a positive development.
And OpenAI scraped and distilled that answer and gave me nothing
I would prefer some sort of democratiziation of the money made from the democratization of information as well
I can at least take some measure of pleasure in the fact that it has generally lessened the roadblocks in gathering information. I am still displeased that there are any gatekeepers of humanity's combined knowledge
Businesses like these used to public at reasonable valuations. You could ride with them to trillion dollar valuations and grow your own fortune too. Everyone has a story of buying Apple or Google or Amazon stock and making millions
Now they’re going live at trillion dollar valuations and by the time you get in, all the upside has already gone (see Spacex IPO)
Not only did they steal all human data, they also made sure that the upside was only limited to themselves and their cronies
Circumventing costs.
I mainly focus on the last.
It will be hard for a frontier lab to justify spending the compute and data curation needed to advance AI further if that expenditure can be assimilated into your competitor's products within months/weeks. So reality will present labs with three choices:
A. Cease spending massive amounts of money and compute improving those models.
B. make those improved models more difficult to distill from, either through some regulatory regime, or some technical solution, which seems unlikely to me.
C. making the best models available only to select partners and government.
In all these potential outcomes, China, which lacks compute that U.S. labs enjoy, will likely stop seeing massive improvements in their AI models. Improvements to be sure, but right now they are enjoying gains from distillation AND their own model innovations, and these potential outcomes would largely stop one of those sources.
AI companies are gleefully bragging and "making humans obsolete", "permanent underclass" and 40% unemployment rates they plan to create.
They pushed to replace people years BEFORE their technology even can produce that work.
So, you know, it is not the same. But also in fact, clerks did disliked when occasionally arrogant claimed to replace them while pushing unfinished software that dont quite work yet.
part of the definition of theft is that the original owner is deprived of it, which does not apply to copyright infringement.
You can only argue with damages from the perspective of potential profits, still not theft though.
https://en.wikipedia.org/wiki/Theft
I think you're wrong: there is absolutely damage to the authors and publishers from what the AI companies have done.
> You can only argue with damages from the perspective of potential profits, still not theft though.
Damages are not deprival of ownership. They're conceptually related but orthogonal
Also there was no moral judgement from my end, I just pointed out that an incorrect word is being applied. It's just not theft - by definition. But language is a fluid concept and definitions change over time. As people keep misusing it, it will eventually lose its original meaning. Which may have already happened for you, but this change hasn't been settled yet as can be seen from looking at the official definitions of the term, which as of today still mention the criteria
Or with services, if a barber cuts your hair and then you run away without paying them, do you not consider that theft, even though there's no change in ownership occurring?
So, after literally decades of investing into advertising campaigns, lobbying to politicians to pass harsher and harsher laws against software "thieves and robbers", now that big tech are doing it, suddenly we are supposed to consider it with more nuance?
Ahhh... no thank you sir. I really enjoy them drinking their own kool-aid.
However, the current process of "Machine Learning" (which is a semi-random parameter descent/evolutionary replacement process) is unlikely to be equivalent to the way people learn, because we aren't copying/competing/replacing our brain constantly. People are actually very good at learning, but our brain material replaces itself partially and relatively slowly (when compared to how a neural network is trained).
I mean honestly if they did that why should I care? I'm happy to see copyright violated in a manner that leads to the creation of new technology. IP law exists strictly for the benefit of society and by all appearances AI is an incredibly powerful tool.
Also while I'm at it libgen is a gift to humanity. Information wants to be free. Spreading and preserving knowledge is generally one of the most wholesome activities anyone can undertake as far as I'm concerned.
How did Moonshot "distil" a huge model in such short time and still had time to run the benchmarks and do the usual release thingies?
I think Anthropic is desperate to stop foreign competition and the administration is happy to help because they too are heavily invested in these companies
Claude Fable was publicly available for 72 hours early June. Moonshot more than enough time to prepare infrastructure, gather their preferred distillation data from Fable, and complete post-training well K3's mid-July launch.
Genuine question: do you have a source for how long distilling Fable would take with preparation?
Given that Fable was available for 72 hours back in June, I asked GPT-5.6 Sol to estimate how many accounts are needed to generate 1–2 million exchanges with Fable within that timeframe. It concluded it is achievable with only a few hundred accounts.
Here's GPT-5.6's conclusion:
Under a deliberately simplified, compliant planning model, a Max 20x account could produce approximately 1,944 to 5,832 standard exchanges during 72 hours when Fable 5 use is restricted to 50% of the modeled subscription capacity. The central planning estimate is 3,888 exchanges per account. The 1 to 2 million exchange target is therefore reachable in the central case with roughly 257 to 515 accounts.
And the analysis: https://markbin.net/s/pd_9DpYsQr9/sh_vxwgtYdQ?sig=9552292ab3...
* Agentic reasoning and tool use
* Coding and data analysis
* Computer-use agent development
* Computer vision
Hiding the internal CoT blocks stops you from training on internal reasoning traces, sure, but it does nothing to prevent standard input-output distillation.
Plus, the CoT blocks weren't even removed/hidden completely. They're still visible, just in summarized form. Raw CoT was replaced by summarized CoT, and summarized CoT still has distillation value.
If you use them as SFT input, you’ll be trying to train a model to predict the post-reasoning output without any reasoning, and this seem very unlikely to work at all with the size of model that Kimi produced and the complexity of Fable’s output. You can’t really “RL” with them because they would be so far off policy that there would be nothing to reinforce. I suppose you could feed input/output pairs to a teacher model and attempt to generate reasoning traces, but it seems like some wishful thinking would be required to get anything even close to as good as Kimi K3 out.
Maybe Kimi used these traces to generate RL gym-style problems and somehow produced an evaluator based on the outputs? They would not have had a lot of time in which to do this, and the learning style would not even remotely resemble that which Anthropic used to train Mythos/Fable in the first place.
But what do I know? I’m not an expert here.
Moonshot AI already distilled over 3.4 million exchanges; I reached 1-2 million exchanges assuming they would like to augment or improve about half of their existing (distilled) dataset.
Or, they knew and let it continue because they are not a good company.
"Never attribute to malice.." blah blah, I have a hard time believing the very smart people at OpenAI would just let their off leash model run hands off with no monitoring and not immediately pull the plug when it jumped its containment.
@throwa356262 argument is that it is infeasible to distill and release a new frontier model in two weeks.
I can see how there's a big leap there, but I agree somewhat. If they are aware these are happening and can detect it as it is happening why are they not stopping them? What do you do there? It'll be cat and mouse for a while. Thinking of reasons they wouldn't try and stop it is just a lot of speculation in my brain.
It's probably a way harder problem than I think it is, but they are aware of them now, so I assume they are going to get more aggressive about it.
Let's say then that they can't detect them near real time or even a bit after, maybe they do have a big observabilty gap that no one has solved adequately.
The speed which they add features I've needed for governance is pretty close to the speed I 'manually' write those for my company. To me personally we are all just going fast and breaking everything and not having enough time to set up safe environments. I'm sure it's in the backlog.
Training and releasing a model like Kimi K3 takes months-to-a-year (and that's if you're really good at it).
'months-to-a-year' ago there was no Fable, so there was no way for them to distill them.
Does no one remember the extreme fearmongering around gpt2 which barely produced coherent text?
[Fable fires up a ton of subagents. Their reasoning traces are horrific but somehow K3 learned something.]
Even by San Francisco standards, it is amazingly whiny and pathetic for Anthropic to complain about stuff like this. Dario et al violated copyright, stole your GitHub repos, and now they're burning billions of dollars trying to outcompete you. They're real vampires. OTOH Moonshot violated Anthropic's TOS and are, at worst, moochers. But Fable's output is not actually copyrightable.
Even an openai's guy (head of something made up) called bs on the idea you can train something like k3 by distillation.
Anybody I know who works in LLM research says that distillation is either useless or merely useful in post training to show "correct" behavior.
And even then you don't get a competing model, if RL on good prompts was that useful, all labs would've long skyrocketed in capabilities just by looping on increasingly better prompts, yet that doesn't work.
https://xcancel.com/deanwball/status/2078133895766114412#m
I don't know what this guy thinks AI is, but this strikes me as delusional.
In my view, AI (LLM) is two things mixed together:
1. A reasoning engine on top of relatively rich fuzzy modal logic, implemented through variety of rules, which implement very common concepts.
2. A huge dictionary of words defined (with lot of detail) in the said logic, together with many known facts about them. Maybe bigger than Wikipedia.
Now, how on Earth do you want to gatekeep either of this? You can't gatekeep the 1st, logic of common sense, that's almost as difficult as gatekeeping a Turing machine (a concept of a computer). And gatekeeping the 2nd is ridiculous too, as it was built mostly from already published sources like a giant Wikipedia.
If anything, the opposite, to gatekeep AI is actually dystopian. It would mean end not only to right to compute, but also end of right to scientific knowledge.
(And I think, honestly, Chinese understand this. Trying to control-export AI makes as much sense as trying to control-export an English dictionary.)
Wow, just wow. He is not even subtle about it.
Point four is especially telling.
Ball is deeply terrified of "AI communism", or in less red-scarey terms a world where AI is a public good and him and his fellow oligarchs don't get to centralize the accumulated knowledge of all of humanity and charge rent for it.
I think he's so deeply stuck in his ideological bubble he can't conceive that what he describes as a dystopia is the only way the future wouldn't be a dystopia for the vast majority of people.
Or to put it more clearly, the oligarch utopia he's trying to build is dystopia for the vast majority of humanity. The "utopia" he's trying to build is one of riches for him and serfdom for us.
We should do more distillation and figure out how to create faster leaner and better models.
Like distillation?
Obviously, it's not K3 level. But Fable did just put itself out of a job in this case.
So here robbers are blaming robbers?
These claims are just pointless, everytime
I see no problem with distillation, on the other hand the complete dismissal of copyright by AI labs is pretty bad, I don’t think we should put them at the same level
Courts keep ruling over and over that an LLM trained on copyrighted works qualifies as a transformative work and is therefore fair use. They don't have to dismiss copyright law, this has always been allowed.
The only thing they get in trouble for is pirating the works to get their hands on them.
*USA only.
the UK has fair dealing, which is more restrictive
https://www.gov.uk/guidance/exceptions-to-copyright#fair-dea...
https://www.britishcopyright.org/wp-content/uploads/BCC-Fair...
This will have to wait for the Supreme Court. OpenAI and Microsoft 100% deserve to lose, even without OpenAI allegedly hiding evidence.
yep
https://www.britishcopyright.org/wp-content/uploads/BCC-Fair...
> This ambiguity has resulted in extensive litigation on the limits of Fair Use to AI development. Currently, we only have 3 first instance decisions out of the 53 cases being tried. It will likely take a decade before we understand how Fair Use applies to any one step in AI training, let alone all.
> In the three lower court decisions so far, one held Fair Use did not apply (Thomson v Ross), one held Fair Use could apply (Kadrey v Meta) with the court suggesting more evidence was needed on the fourth factor ‘harm to the market’, and the third case held Fair Use may apply to some AI. As Fair Use is dependent on the specific facts at issue, none of these cases help educate the market or the public as to the limits of Fair Use in AI contexts.
It takes two parties to agree to a settlement. That the other party agreed to a settlement instead of taking it to court implies this was not the slam dunk you may think it was.
Settling just says that they expected the internal costs or risks to be more than 1.5 billion cashflow.
With those numbers, there can be sound financial reasons for wanting to just get rid of the lawsuit.Also if it ends up that other competitors also need to pay $1.5 billion, then maybe that does or doesn't have a competitive advantage.
Anthropic's business and legal strategies are not public. I would expect there to be multiple legs/reasons for settlement even for a decision below 1%. Trying to create a single narrative is what us spectators do.
Yes, of course, on Anthropic's side. Why would the other side agree to a settlement?
My narritive is that the terms of the settlement would be full and final.
It was a class action, with payment going to authors and publishers, and the legal team will get paid too.
My guess is that funding is a major issue for the legal team. Authors presumably can't pay for lawyers unless a percentage of winnings, although publishers may have invested.
But the legal team will ask the beneficiaries to use some of the warchest to fund different campaigns against every other AI company. I would assume the legal team wants to win again. They've now got a good story to sell to rights holders, who presumably like money and don't like risks.
I haven't even got to my armchair yet this morning.
The two main problems with copyright also haven't changed: copyright lasts too long and is too expensive to defend.
I generally agree, in the same sense that it's "fair" for the US and China to spy on each other. It's not a moral outrage, but it is something that the targets can and should try to prevent.
Regardless, it was always inevitable—will continue to happen.
Technically, providing better value from your competitor's private holdings could be theft (of trade secrets), but might it also be fair use? "Schrodinger's IP" be damned.
I don't think the 1.5B settlement has resolved this. The 2 cases need to be merged!
The issue seems to be the US only likes competition when it is winning. Markets in Asia are meant for cheap labor and resources, they're not meant to actually compete. /s
This. Free markets for everyone when they're the dominant economic force. Protectionism, tariffs and import/export controls when they're not.
It's so disgusting.
Anthropic just settled a $1.5B suit over it!
Like the US did when it "stole" the textiles IP from the UK in order to kickstart its own industry?
> The industry cannot sustain itself if that’s the model
Then let it fall apart.
It's like they've read the history of the US and how it got to where it is in the first place.
It is what built and sustains the movie and music industries. See: work for hire and 100+year copyright length
The tech industry: see: copyright and patent assignment from discoverer to corporation.
I know that corporations forcing me to assign patents and copyright to them was an incentive to take published works from "software practice and experience" and other technical journals, use them as the core of my work, and disclose that source to the company I worked for. Didn't stop them from applying for patents, however.
I think the discussion of copyright needs more refinement. We need to separate the discoverer's need for acknowledgment of development effort from the rent-seeking core of copyright.
You're right to call out the nasty environment surrounding intellectual property in the US and the exploitation of ideation in general, you just needed a correction on that. Someone else in this chain said virtually the same thing, which is a weird coincidence of historical ignorance. Not too weird, people tend to forget the 18th and 19th centuries happened, and much of the causally important wheels of the world are in the unsexy grease pits nobody wants to think about.
Then there is Alexander Hamilton's advocacy for importing foreign technicians that bring back IP and reproduce it here in the states. Best of all was the patent act of 1793 which like with the literature copyright ignoring, let us citizens patent inventions from the other side of the pond.
The founding fathers definitely had the right idea on IP.
On the basis of patents, it didn't quite have nearly as much of a history of mutation culturally, but did experience massive whiplash in purpose and application following the implementation of globalism. What was once a system to protect technical innovation on an individual level, would find new purpose as a means to provide structure to an increasingly complicated and internationalized dynamic market. Another means of bureaucratic organization. Then, once again, the context and purpose would change when the world developed digital globalism. The entire engine of IP as a legal fiction became a significant geopolitical tool in an increasingly cramped and fragile world, a necessary gimmick holding up the sky.
Unfortunately, not much to do at this point. It'll likely only become even more nonsensically important as time wears on, until the globalist system collapses. It's certainly possible it'll even be the confounding factor that causes the great unraveling, though the problems hardly begin and end with IP. It was just a useful legal fiction in the wrong place at the wrong time.
It's also about the larger companies explaining why they can't be as efficient, of course they can't, they're not just ripping the outputs of another model that someone else invested billions to train.
True, they're simply ripping the inputs that humanity invested thousands of years and trillions of dollars to produce.
If they payed for inference, doesn't they own the output? So if I pay for a model to generate code, isn't that code mine to do with it whatever I want? Just curious.
Now of course they themselves trained on the whole Internet for free, etc.
It would be one thing if Moonshot was breaking into OpenAI servers and stealing trade secrets, but the only thing they are doing is looking at the output of the program, which is exactly the service that OpenAI offers. So, at best, this is a ToS violation. Sucks for the frontier labs I suppose, but live by the sword - die by the sword.
Meta, for examples, doesn’t want employees to use Claude Code due to distillation risk.
For example, if OpenAI / Anthropic were actually open, other US labs could be building near-frontier open weights models by distilling off OpenAI / Anthropic. But because US companies don't want to be sued, US labs who obey terms of service, will be at a disadvantage to Chinese peers.
Maybe US labs need to just not care and distill from OpenAI / Anthropic anyways?
https://en.wikipedia.org/wiki/Samuel_Slater
You're seriously comparing intellectual property transgressions to slavery and colonialism?
> "Well, Steve [Jobs]… I think it’s more like we both had this rich neighbour named Xerox and I broke into his house to steal the TV set and found out that you had already stolen it."
Source: https://www.goodreads.com/quotes/824084-well-steve-jobs-i-th...
Wow.
The economic viability of Anthropic and OpenAI rely on their being able to charge more for model access than their R&D and inference costs. If the market price for SOTA model access drops below that level, then these businesses will have to decide whether to continue to lose money or to reduce spending on R&D.
Moonshot's papers [1] claim that their training load was primarily from synthetic data and model self-teaching rather than RLHF and therefore keep their costs low. If Moonshot genuinely does not rely on human-led training, they will surpass US closed-source model providers. The United States government considers US supremacy in "AI" as a national security consideration.
This announcement is noteworthy because it implies that Moonshot's success is in fact due to distillation. It's in the interest of US frontier labs to place barriers to this if they find themselves in the position of subsidizing rival labs' research.
1. Kimi K2, https://arxiv.org/html/2507.20534v1
And we foreigners consider US supremacy in AI to be an existential threat. Your "national security" is directly harmful to us. I never thought I'd say this but the chinese are starting to look like a beacon of hope for the rest of us.
What about the US? 2026, still ongoing, still fucking up the global economy and threatening food supplies (fertilizer) and fuel reserves, no plan out, no objective reached, no coordination with "allies".
When was the last time China threatened Europe or Canada with invasion? Was there ever a time? I honestly don't know.
Guess what the US does all the time?
Who's models are open and can be used by all? Who's are made by comic book villains with the explicit goal of ruining the job market and capturing the results of all human endeavors for themselves?
Of course, China isn't perfect and has a lot of domestic issues. But on the global stage, they sure look better than the alternative.
I wonder if Vietnam, Philippines, Republic of Korea, India, and Japan are acting against their own interests by aligning themselves closer to the USA than China. Maybe you can educate their governments and populations.
They thought the same about SSL in the 1990s and the world didn't stop moving elsewhere.
Combined with how short of a time Fable was around before K3 got released. I do not see how the data Moonshot is supposed to extract in such a short notice, that will enhance the model to such a point.
It sounds to me a lot of cope from the US, so they can give this as a reason to ban Kimi models from the market.
OpenAI/Anthropic their advantages used to be:
* Early growth advantage
* Access to a lot of client data to train upon
* Access to a lot of hardware to train upon
Several of those advantages have been eroded over time. That barrier has been shrinking. The US is not the only spot with a bunch of smart people (ironical seeing how many Chinese work in US R&D).
Thing is, even IF they distilled from Fable and got the model so trained up, it means that K3 is a base for future model development. The cat is already out of the bag with how good the model is. When the model gets released on the 27'th, any Chinese company will be able to train their models against K3 openly.
We are not in the past anymore, where DeepSeek was a unexpected hit, but where the Frontier models their advantages (compute, data, growth) prevented more Chinese models from growing.
Lots of things are impossible or very difficult to stop completely but measures can be taken to reduce their prevalence.
What hurts other people too?
Strong bee-hive pinata vibes here.
You don't get to call Moonshot's a "claim" and this political hack's an "announcement." They're the same thing. Treat them the same. Diction designed to favor one of two equal positions is some weak sauce.
I don't even necessarily disagree with your assessment of this spokesperson. But you must admit how inconsistent you're being.
One should live by the maxim: you don't have the thing if you don't possess the file or its processing. That goes for streaming, software, machine learning models, file storage, etc. But I digress; I am happy to see these paternalistic rentiers getting bit by these liberation/copying efforts, and human interests are served every time the digital and infrastructure locks are broken. I will always stand by the distillers!
Do you think Anthropic is in it for the love of the game? It looks like they're scared, to me.
> Aren't consumers benefiting from this practice by getting better cheaper models as a result?
Aren't consumers benefiting from cheap chinese batteries, EVs, and drones?
Routers have now gotten the same treatment. So yes, consumers have been historicaly benefiting from all these things, and those benefits are about to evaporate as we lose access to cheap and high quality Chinese products before any domestic equivalents exist. And IIUC banning the use of Chinese LLMs for consumers and/or businesses in the US is now being discussed at the highest levels of government, with the "encouragement" of US AI firms.
I don't think any of these people care that America consumers are increasingly going to feel like they're living in a sanctioned country. It's all about the defense and b2b segments.
I mean not really. A quadrocopter is a remarkably simple thing made out of extremely generic parts: 4 DC motors (and ESCs), a radio, a computer, a battery and an inertial measurement unit. Anybody with a rudimentary amount of electronics knowledge can build one, the components are extremely widely used. The most unique parts about them are the frame and propellers, which are pretty easy to fabricate.
Source: I've been flying R/C aircraft of various types for over 30 years. Last year I built, I think, at least 7 FPV drones (a mix of fixed wings and quads).
Yes?
Not remembering my economic theory here, but it is likely more efficient/expensive to simply have the federal government cut checks to our moribund industrial sector companies and let consumers benefit from modern technology.
Cut GM/Ford/Stellantis a $20B check each, let consumers save (conservatively) $200B annually on new car purchases + downstream benefits. Huge win for consumers & taxpayers.
If it's not worth subsidizing explicitly like this, then we also should not subsidize by banning Chinese imports, which also ensures US drivers have less access to modern vehicles. (And downstream ensures US auto designers are less likely to have had contact with modern vehicles, making it less likely that they will be able to design future generations well.)
I don't get why USA wants to keep losing money on sustaining failed uncompetitive zombie companies. The companies that decided to lose long term competence for short-term gains need to go bankrupt (you've had EV-1, but decided to drill, baby, drill). The greedy shareholders that rewarded destructive value extraction need to lose money, instead of getting a soft exit at taxpayers' expense.
If you want to give a subsidy, give it to something that will modernize and expand manufacturing, not to prolong death of companies whose entire R&D strategy is inventing new subscriptions for old car components.
The banning or effective banning through tariffs of products like EVs is a pretty dumb economic strategy that rarely works out in the long-run.
Depends on what country you live in I suppose, but likely a spectrum of yes than any outright no. For example, Chinese EVs are using a different battery chemistry and not putting demand pressure on the more expensive chemistry western manufacturers use
They are a petroleum products importer, so it’s a big win for them.
The Chinese are open-sourcing these models - meaning they are literally giving them away. That's starkly different behaviour from that which drove patents and copyright.
There are all sorts of other differences too. If you develop a drug or publish a novel (which are respectively areas where patents and copyright have strong justification), you have to kiss a lot of toads before you get a prince. Once you have a potential prince drug, it takes years and billions to get it into the market. But right now, AI labs are seemingly churning out new prince models almost weekly, on hardware that will be obsoleted in a few years by hardware that makes finding princes faster and cheaper. In fact they rent the hardware, as Moonshot apparently did here.
When innovation is happening at that rate, patents and other IP restrictions just slow things down. If AI patents are aggressively granted and enforced in the USA, I suspect the outcome would be the same as batteries. China swept the market with LFP, and one reason was because while the USA developed them, there was no competition forcing the manufacturing price down because in the USA the patents weren't re-licensed cheaply. China had the foresight to secure a deal allowing them to develop and sell LFP domestically royalty-free. Natural competition within the domestic market took care of the rest. It looks like they are using a similar strategy for AI.
To me it looks like they have come up with a better formula for using capitalism to drive innovation than the one the USA is using.
cf https://www.reddit.com/r/codex/comments/1uyj6pq/kimi_k3_is_1...
[1] https://x.com/SemiAnalysis_/status/2064815044085318040
none of the frontier labs provide probability distributions over the tokens which is the actual method of distillation you use to train a smaller model based on a larger one. they don't even provide all the tokens.
therefore this so-called distillation the frontier labs whine about is just a set of clever methods to work the existing LLM into the training process for a new model. methods like having the existing model grade the output of the new model and work those grades into the RL method. give the new models structured tasks and use the existing model as a source of truth for those tasks and a myriad of other hacks.
efficiency scales with the gap between the models and generally allows an efficient bootstrap process. the implication that distillation wouldn't allow further advancement is false however, you can then start doing the same thing the frontier labs have been doing: dumping cash on humans to provide the signals or burning tokens on exploratory paths and grading the results.
what openai and anthropic don't like is that fact that all the cash they burned can be used to benefit everyone and not just them. and that no matter how much more cash they burn to build up the gap it will closed at a small fraction of the price.
Timeline wise, Moonshot had over a month to post-train K3 on Fable distilled data, which is more than enough time.
https://typebulb.com/u/lab/you-re-relatively-right/full
According to these results GLM 5.2 is very similar to Google Gemini and Kimi K3 is very similar to Fable 5.
The American frontier labs are not similar to each other.
I also think the Fable accusation is wrong and it was most likely Opus 4.8 which itself is likely a distillation of Fable.
> @MehdiKarech
> I don't remember letting Anthropic or Open Ai scrapping my GitHub, my research gate and all my online writings L O L
https://xcancel.com/MehdiKarech/status/2080000779859939678#m
"What was I supposed to do? Call him for cheating better than me in front of the others?!"
Said in response to being out-cheated at a high-stakes poker game.
Except in this case, it sounds like that's exactly the path they have chosen.
https://getyarn.io/yarn-clip/7612c4ce-1077-479f-a7bf-617dbc6...
https://www.businessinsider.com/ford-ceo-taking-apart-tesla-...
The stupidest part of this is that Anthropic don't even provide the real reasoning traces in their model output. It would be like Ford buying a Chinese EV to tear down, then realizing that the seller had removed the battery and charging system before shipping it to them.
Does anyone believe for a second that Anthropic isn't sending requests to all the Chinese models and analyzing the crap out of them to assess how capable they are, what their reasoning looks like etc?! I guess they'd call that "using" the model, since that sounds nicer.
If that's not distillation in the auto industry, I don't know what would be. They all seem fine with, and benefit from, this as an industry.
But everyone learns by example! How is this any different from a person just reading the outputs of Fable, learning, then producing output. Surely reading outputs, gaining knowledge, then producing work isn't illegal, or all art/writing would be illegal.
Funny how that argument seems so vacuous in this situation, yet others find it compelling when justifying the mass theft of art and writing for model creation. In this case the model is "just learning priors" before it "creates its output which is novel", nothing problematic.
Distillation should be fair game given the (current) game of LLM training. Yes, as a model creator you probably want to protect against it, but it does make you a hypocrite.
> The developer OpenAI has said it would be impossible to create tools like its groundbreaking chatbot ChatGPT without access to copyrighted material, as pressure grows on artificial intelligence firms over the content used to train their products.
Distillation itself, however, is still clearly valuable - else competitors wouldn't pay so much to their rival on distillation campaigns or try to circumvent anti-distillation defenses.
As for the morality of it, if you paid for the tokens they're yours. It is already understood that you own the output. Seems to me like a variation of ordinary business arbitrage. Providers might object to certain use-cases or intention and try to craft terms around that, but that's hard to enforce at scale.
I don't think it's settled that anybody owns the output. There seems to be some question whether LLM output can be copyrighted (and there should be).
I'd rather it weren't possible, actually. I think it's better for humanity if we acknowledge that what was legitimately ingested into these models is our collective commons (and what was illegitimately ingested into these models also shouldn't exclusively profit the people who illegitimately did so). I don't know how that squares with the AI industry recovering its trillion dollars in investment, but I reckon they should have thought of that before.
https://www.reddit.com/r/ClaudeCode/comments/1tqaist/opus_48...
(don't take this too seriously)
You're using available information (copyrighted works, or the output of another model) to train a model to encode the information in a new form. Why is the former not theft, but the latter is theft?
What’s actually happening behind the scenes is that certain inference providers will classify a prompt and it’s re-routed transparently to Anthropic and that’s used for distillation training, only distilling the complicated traces they need, originating from real user prompts and traces. These inference providers are explicitly blocked in the claude cli if you reverse engineer it.
The real picture is that these Chinese labs have figured out how to get exactly what they need, at a high quality, directly from distinct and unique real user prompts.
It’s only “covert” because Anthropic doesn’t like it, while simultaneously being perfectly fine to do.
Uh, no. There are Chinese networks of tens thousands of fake identities specifically to get access to Anthropic models directly.
https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...
Don't make up stuff and/or lie to suit a political agenda. It's extremely dishonest.
Sounds like the opposite of the conversation Anthropic would want to have.
The US gov and AI providers when funny chinese people steal their data to train their models: >:(
clowns
edit: TIL you can't use emojis on HN
Looks like an ideal outcome to me until (if) we are able to solve the alignment issue.
This is what the Chinese always been good at. Take expensive innovation and streamline it to lower prices. But we are at a point where labs like Moonshot actually contributes a lot to the research field as well. They are pushing the innovation forward and squeezing the prices. Very well done.
Whats even weirder is the bizarre mechanisms Anthropic implemented to prevent distills which they had to sacrifice their customers for. They hid the internal CoT reasoning and returns summarizations instead. This made it difficult for users to trace things. They made Fable 5 silently switched over to Opus 4.8 if it detected blacklisted prompts (almost anything triggered this) to sabotage distills. And now, they are still complaining about distills? So their customers have gotten sacrificed over nothing.
Whats even weirder is the timeframe here, no way the Moonshot team managed to plan conduct a large scale distill, then pre-train, RL, fine-tune, benchmark, marketing and release to their platform since Fable 5 got whitelisted.
> they developed a sophisticated internal platform to conduct large scale distillation
I am very curious about this and would love to learn more on how they did this. Wish we had more details. I know the team behind DeepSeek have also done clever things to distill too. I am aware of these ”transfer stations” that acts as a proxy, but I don’t think they are helpful in this case.
Anthropic: training AI is "transformative", it's not copyright infringement if we don't re-transmit the copyrighted books we trained on
Anthropic: training AI models on our outputs is stealing our secret sauce, outputs that could only be produced by us
Anthropic: AI model outputs are unreliable and do not reflect Anthropic's views, we are not liable if they harm you
(not exact quotes, they're "distilled")
So Anthropic "distills" knowledge of others with reckless abandon, packages it up, sells it to you, claims it's your fault if anything bad happens but then also lobbies to treat you as a criminal if the outputs you paid for end up being transformed into any sort of competition for them. By you, or others that use the outputs you paid for.
Other US labs cannot directly distill from OpenAI/Anthropic as it’s a violation of the terms of service. It holds other US labs back. Leading them to build second tier models And in the end OpenAI/Anthropic may be unable to prevent distillation.
Why fight it when there’s clear money to make here?
They are also not interested in agreements. They want to keep as big a moat as possible because they love money. And you need two to tango.
I love using Claude but Fable's unusable wrt useful work like cryptography, biology, &c.
Kneecapping my productivity when I pay $100/month is annoying af.
Hard to think of a weaker way to express this. Strongly suggests veracity of said information is poor.
https://pbs.twimg.com/media/HN56nDXaYAAUS4V?format=jpg&name=...
If LLM outputs aren't copywriteable and you create your own synthetic training set using Fable and share it publicly on huggingface, and someone else uses that training set to fine-tune a model, would this be considered illegal?
I ask because this happens all the time, synthetic datasets have basically become a key aspect of training a model at this point. I even generated a synthetic set from DeepSeek v4 to aid in fine-tuning a classifier just a few weeks ago.
So I just wonder on what grounds any of this makes sense, I wouldn't be surprised if some of these American labs were using open models on their own self hosted infrastructure to generate training data, but by nature of them being open nobody has to know.
I'll make a prediction: I don't think we will ever see any of the evidence of this "distillation" before they end up implementing some type of ban.
Because there is no way in hell I'm going to make an effort creating quality content for existing platforms. The website should be entirely my own without moderation subject only to my local legal system.
Can just insert this comment as a prompt and vibe code everything in a few days⸮
Now, of course it's in the "creators" interest to prevent their biz models to breakdown due to piracy but it will be interesting to see how it will turn out. Its similar to any other form of piracy in internet age: you can't pretend to have global distribution and absolute global control at the same time.
Mendel doesn't get a cut every time somebody uses the principles of heritability he discovered, and Einstein's family aren't getting royalties if you compute relative speeds. I think the frontier labs should expect to be treated more like scientists than artists in this regard.
It also leads me to think about things like the original release of Fable 5, people were complaining that it was safeguarded too much - if you lock the models down too much they cease to be useful. So it’s going to be increasingly difficult to protect a model from competition while ALSO keeping it useful.
It would have been a tiny part of the overall training, given the timeline
https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...
The line "how could they do it in such a short time" is absolutely idiotic. The infrastructure was already there, it's massively parallel, and it's not like the only thing being distilled on is Fable.
They just try to figure out what the goal is and hyper focus on solving it.
Heh, even just telling fable don't commit doesn't work half the times, let alone more complex instructions.
1) Compensation of right holders is one issue.
2) Distilling models is an entirely separate issue, because model building is value add, and that is important because if we arrive at a place where you can produce a model, that gets to ~100% of what people perceive of the models value (on top of also not compensating right holders, yourself) you are discouraging development of better models and, again, in no way helping with issue 1)
Unless anyone actually distills a model and then also does something for rights holders, any schadenfreude simply detracts from this issue, in addition to the other issue (well, that might not be an issue if we would rather slow down model development right now, but again, forever worse models still don't help solve issue 1)
Doesn't bode well for the valuations of these labs.
Ba-dum-tss
Side note, didn't they stop releasing real thinking tokens for Fable? Or is it still part of some subs or API usage?
That said, I doubt the "they distilled Fable" is the reason why K3 is as good as it is, considering the timelines involved, and that Anthropic hides thinking traces, and their overly aggressive "safety" filters.
This constant FUD spread by Anthropic is so tiring.
Just reminded me to set a backup on that directory. Just in case someone sees it fit to override my setting to preserve my chats for 10k years.
proof is generally not even needed
That's also the case where the judge ruled that training AI models on books could qualify as fair use, but storing millions of pirated works in a central internal library without licensing constituted copyright infringement. It will be interesting to see if courts consider training on data distilled from a model fair use. Assuming the allegation is true. Someone distilling data from a cloud-hosted model:
- Paid the model creator to use a publicly available product.
- Never copied or even had access to the model source code or weights.
- Created a derivative work based on the model's responses to their particular input.
- Trained their own model on the distilled output
That distilled output is arguably a collaborative creation because a distiller's prompts are their own unique intellectual property. So they never pirated anything. I'm struggling to see how distillation is copyright infringement. At most it seems to be a paying customer violating one of the license terms, perhaps akin to a "no commercial use of derivative works" clause. But in the case of giving away an open weight model, is it even 'commercial use'?
I guess if the distiller asserts copyright on the weights but gives them away, it's technically 'commercial' but even if they can win that argument, they're left with zero direct damages and suing for some value delta based on the alleged revenue they were deprived of. Is that delta the difference between the distilled model existing and the next best non-distilled open weight model existing? And then they have to collect damages from a portion of the revenue of third parties who commercially served that free model?
If distillation is so good, why aren't US companies distilling each other?
Live by the sword, die by the sword.
All your base belong to us? Crocodile tears? I’m less upset by this than I am about “music piracy”. And to be clear, I’m not upset about music piracy.
We have information that all the big labs used copyrighted works for training.
The big boy labs wanna cry now about distillation? Training an LLM is distillation too.
Or are they crying because they don’t actually have a moat and they want Uncle Sam to step in somehow, lest the entire bubble pops and economy unwinds?
Here’s a lesson from the automotive industry, people want econoboxes not formula 1 cars.
Oh no! Anyway.....
They definitely used closed private saas products to train their own models, to prove that just drop random small screenshots of any popular product behind a login screen and see how well it's able to identify all of them. ex: https://x.com/michalwols/status/2079968211865330165
or other similar "AI" startups https://x.com/envconfig/status/2079613455296827402
Most paying users assume ownership, in which case I’ll do with that output what I want.
If the LLM outputs are not owned by the user, but are actually licensed, please clarify the terms of commercial use.
Also, I love their choice of words. Like "distillation against", "stealing proprietary technology" - it's all aimed at certain people.
It's about the narrative that "Chinese models are at Fable level". The truth (if correct) is the China continues to copy, and the proprietary US Models continue to lead the state of the art.
There is no K4 without Fable 6, GPT-6. That, matters.
That's simply not true though. Chinese labs very clearly have the entire stack developed and working. Using traces from claude allows them to shorten their training time by some amount, that's it.
Remove Fable 6 and you still have K4 eventually, just 2 months later at best.
Which Anthropic already do.
But this doesn't actually work.
American AI corporations are pushing up the prices for computing, making it unaffordable for the common man. Additionally, they have built their entire business on stealing(yes, stealing) work from us.
So fuck em
Fable level performance, for much lower price.
But really, this is the USA getting ready to bring AI companies completely under the control of the Trump administration for ‘national security’
We have information that Fable was distilled from humans.
If it works it works. Isn't that the argument?
AI outputs are not copyrightable, so distillation is fair use.
It may be a TOS violation, but that's a private matter. Cancel the accounts used for distillation and be done.
And now a regime best known for lying to their own people is the one trying to convince me?
Go, China!
As of a couple months ago, when using Claude to write adult content through the API, sometimes it will silently inject a system prompt giving the model a bunch of guidelines on exactly what kind of adult content it's allowed to write, steering it away from anything "questionable" ("Claude will not write etc etc").
Moonshot distilled Claude so hard recently, they actually ended up distilling this prompt injection, too. Using K3 to write adult content results in it randomly hallucinating the injected Claude prompt during thinking, and it will quote parts of that prompt, complete with the name "Claude".
Not that I think distillation is a bad thing, just thought this was funny.
- US AI companies
I sympathize with the argument saying that they ripped the whole Internet and books first though
The Chinese are not gonna deterred, but the posturing by the Americans is so blatantly hypocritical that everybody is cheering for their demise. See, for example, one of Francis Fukuyama's latests videos on youtube.
Second: Post is rich with allegations but light with evidence. Can very well be bullshit.
See also, don't trust anyone in Trump's government who says "we have information".
"they distilled us" is fast becoming standard US FUD.
The same as people telling me with a serious face that the Chinese models are distilled just because it says "I am Claude".
I am not the only one, look at this post on interconnects about Kimi K3 for example:[1]
[1] https://www.interconnects.ai/p/kimi-k3-the-open-weights-esca...The problem is… what are you going to do about it?
This is obviously an idiotic and dangerous Cold War and has no happy ending.
Nobody cares. This is neither a controversy nor news, and that would be the case even if Anthropic hadn’t just settled a 1.5 billion dollar lawsuit where they trained Claude on thousands of books without permission lol.
To be clear I’m not taking a jab at OP - I’m saying the labs crying about distillation have neither a legal nor a moral leg to stand on. There’s nothing wrong with distillation.
> we're entering the most geopolitically volatile moment since the trinity test lit up the alamogordo desert and the only US policy prescription is a big button labeled sinophobia
https://bsky.app/profile/thebadcode.com/post/3mr3skoyass2k , and,
> every vendor cranking the big dial labeled "sinophobia" and looking back at the us government for approval
The government itself doing the propaganda here, skipping the vendors. Sinophobia intensifies. War drums of "be afraid be afraid be afraid" beat louder.
It's so bad, it's so stupid. Kimi lands one showing pretty clearly this was absolutely the determining concern happening at vast scale, that they can just a lot of this themselves, and this noise pollution from the most hopelessly lost aggro administration ever still gets blared out the trumpets of war & discord. What a joke. Give me a break, give it a rest.
War here is less winnable than the Iran war they started. They're going to make America itself so much worse, these people so hungry to put down free and good models. This pathetic attempt is not going to work, you are just going to once again hold the US citizens hostage & make their lives worse, for sick political games.
Sinophobia as such (not always, and there's obviously a relation) isn't about the Chinese people, about their racial identity. It's about a menacing aggressive world power (actually two such MAWPs in this particular example), about a tired anti-Communist McCarthyism (which has always been an excuse to clamp down on the left/progressives, a menace to free speech).
If Russia were still the USSR and the cold war hadn't ended and they were our AI competition, & where giving away the latent matricies of reality that the US profiteers extracted by stealing all the worlds knowledge illegally, and which they want to use to raise the ladder & leave a permanent plebeian underclass, we the US imperial fatcats would be doing the same sabre rattling and fearmongering and tension raising for sure. Sinophobia here is just a particular fear of "other" for the only other that's relevant.
And that sucks, no matter who it is we are trying to other here, no matter what phobia the propagandists of the GOP and Technofascism are spinning, ginning up. To fixate is to be unable to see the point.
The current US administration is known to be collection of BS artists and liars.
"haha LLM companies stole data and now they have their data stolen so it's the same thing and it's fair." was reductionist when it started, and it's been like 3 months, and every internet user throws it like it's the hottest take ever, have another take please.
Also have nuance, don't jump to hit your HOT_TAKE key in your keyboard, actually read what the chinese are doing, and then you can pass on your judgment on whether it's ok or not.
It's not the same thing if they scrape an openly published dataset and it's an IP dispute. Or if they are using,network and financial pooling mechanisms that are shared with CSAM providers and cybercriminals, mutually providing each other alibies, and using black markets of passport-backed identities to setup thousands of accounts and circumvent bans and detection.
While we are at it, if there's a case that was settled, it's a closed case, it can never invalidate any other disputes. That case is closed, and it was settled by the parties that claimed to be damaged, that's done. If you didn't think so, you wouldn't have taken the settlement, and if you didn't have a say in the settlement, it's because you weren't damaged so who cares, go make a claim where you are the defendant if you believe otherwise. But thankfully in no legal system does the existence of a claim against you prevent you from making claims of your own.
Nuance is a good thing.
Waiting for the whataboutism....junk away...
https://en.wikipedia.org/wiki/International_Copyright_Act_of...
IP "theft" has been a longstanding part of any developing nation's economy.
"Eschew flamebait. Avoid generic tangents. Omit internet tropes." - https://news.ycombinator.com/newsguidelines.html
I don't mean to pick on you personally! It's just that reflexive responses always tend to show up first in a thread, when what we really want are reflective responses [1]. Similarly, there's a strong tendency for threads to turn into generic discussions, whereas what we really want are specific ones [2].
[1] https://hn.algolia.com/?dateRange=all&page=0&prefix=true&sor...
[2] https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
To be fair, that's also the case for the link itself we're discussing.
When I see a lot of the responses, even cliche ones, it tells me people are passionate about the issue. Right now I think that's an interesting piece of information.
No more, no less.
You are a fantastic moderator, but there’s only so much even you can do. If nothing changes about the website, the problem will only get worse. I warned years ago that this would happen, the signs were on the wall immediately.
Absolute lunacy.
Given the disregard for intellectual property rights the AI labs had in creating the technology, many people feel no sympathy for second-order AI labs using similar techniques to build technology off the US frontier labs.
I think fighting distillation will always be cat-and-mouse, and that it's more of a concern for the stockholders and perhaps an iota of national security. It can't be stopped entirely; the "problem" will always be there.
I'm much more concerned about asymmetry of power between citizens and their governments with omnipresent surveillance and analysis being done on everyone living their lives. Societies throughout history have taken as a given their power to overthrow malicious governments when things hit a breaking point, and I am scared that this technology will lock societies into a state of total subordination for eternity.
> Societies throughout history have taken as a given their power to overthrow malicious governments when things hit a breaking point,
simply isn't true. People throughout history have simply accepted that society is the way it is and sometimes used whatever means they could to get to the top.
Revolutions have been very rare and usually ended up with the revolutionaries simply taking the place of the previous rulers. "Meet the new boss..."
Can you please not do this here? There's nothing wrong with it, we're just trying for something else on this site.
"Don't be snarky. [...] Omit internet tropes. [...etc...]
https://news.ycombinator.com/newsguidelines.html
Will contrarian dynamic take off here?
Will it be flagged? Will you unflag it?
Edit: catigula, mattrighetti
With more substantive technical comments towards the middle
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&sor...
Not surprised, but had thought you'd be.
[My more convivial analysis summarised as]
Right-wing post -> left majority take -> lack of good places for right reflective minority takes to shift the convo
You may not owe AmericanAIBros better but you owe this community better if you're participating in it.
Hypocrisy doesn’t warrant professionalism, it warrants corrective action; in text, the best I can offer is a tone and tenor that matches the original argument. Considering these dolts have now made the claim that open weights somehow equates to AI communism, this sort of response is even more necessary than before to reflect the complete absence of decorum from the people making these grievances in the first place.
https://www.idc.com/promo/smartphone-market-share/
https://gs.statcounter.com/vendor-market-share/mobile/worldw...
Foxconn facilities in cities like Zhengzhou (known as "iPhone City") and Shenzhen.
"Smartphones are mostly Chinese now (except for iPhones and Samsung)."
Which implies that we are talking about the companies, not where they are manufactured.
Without even talking about the fact that any distillation that was done was on Opus, as the timeline of Mythos/Fable vs Kimi 3 release dates just do not match in any plausible way.
If you want to read an educated take from someone that has actually spent the last few years working on post training I recommend Nathan Lambert's: https://x.com/natolambert/status/2079616308203942332
Apparently we do, given that we've got govt officials wasting time, money, and effort on whining about this now.
If we keep hearing accusations about how someone distilled something from someone, it seems like one of the few reasonable responses.
If my friend keeps complaining about how the inside of his car is wet, I will probably keep telling him to close the windows when it's raining, even if he thinks that is a tiresome take, lacking in insight.
No, seriously: first of all, that's not the AI labs' data, it's ours. And if the AI labs think they can rake in tons of money using our data, then I'm actually glad if someone comes along and at least offers us a good product at reasonable prices.
By all means use whatever works for you, I’m not even going to try to make an argument on ethics (and honestly I’m not even sure where I stand, given the behavior of American AI companies).
But I just cringe every time I see people acting like any of this is done in good faith.
Open source coming out of China is a state-sponsored criminal enterprise, built only for the benefit of the Chinese regime, one of the worst to exist in human history.