Rendered at 23:14:37 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
ballon_monkey 3 days ago [-]
The 2 things people need to remember:
1) China can (and does) use the models to influence the west. They train in false information about Taiwan and Hong Kong. Or pretend like history is in favor of China.
2) Ignoring the models containing false information, they are incredible. But you should be scared of running inference via the model creators directly. If you think your data is safe compared to running it via model providers in the US ( either frontier or model hosts like fireworks.ai ) then please let me know your bank details so I can poke around.
Why should I trust a US company more than a Chinese one?
theptip 2 days ago [-]
Your exhibit is an obvious anti-distillation technique. You can expect far worse from less-regulated companies.
throw1234567891 1 days ago [-]
Neither Anthropic nor OpenAI are regulated in the "regulated business" meaning.
theptip 22 hours ago [-]
I count being incorporated in Delaware as a high baseline level of regulation vs. random Chinese API providers.
throw1234567891 5 hours ago [-]
From where I sit, it could be incorporated in Mauritius.
bamboozled 2 days ago [-]
Agree, China believes in Climate change and are at least taking steps to address it. Personally this is making me trust them more than the USA because, facts.
I mean lesser of two evils thinking, if one is intentionally leading us towards climate disaster, while the other isn't then yeah. What else can be said? Should I trust the authoritarian country who believes in engineering and science, or the one that doesn't?
WarmWash 2 days ago [-]
>Agree, China believes in Climate change and are at least taking steps to address it.
China brought 80GW of new coal power online last year. The US added 0, and plans to add 3GW next year (we all know why).
China doesn't care about climate change, they care about energy independence, and conveniently have very little natural fossil fuels besides coal. Which they heavily mine and utilize.
DanielHB 2 days ago [-]
The new coal plants are not planned to run a full capacity (or at all) they are meant for backup in case of environmental disaster and geopolitics (war, sanctions, tariffs, etc). So the actual amount of coal being burned is still going down.
Although it is highly suspicious why they are building so much unused capacity, it is as if they expect something to happen soon.
speedstyle 2 days ago [-]
They commissioned 78 GW of new capacity in 2025, but actual coal-fired generation fell slightly (by 90 TWh, ie 10 GW). It might rise again, but it's expected to peak by 2027–28. And of course they added 440 GW of renewable capacity, 64% of global installations.
I certainly agree there's other motives (renewables even net them significant trade, like EVs) but it's not as though 'caring about' climate impacts requires selfless virtue, they stand to lose $trillions/year like everyone else.
throwaway2037 2 days ago [-]
Woah, I am blown away by these numbers. I did some research using ChatGPT. A coal power plant called Guoneng Qingyuan Power Station is 4GW in total size. (It added 2GW of capacity last year.) Normally, 1GW is considered a large coal power plant by itself. This is a massive plant!
I asked ChatGPT how many coal train cars are required per year to fuel this 4GW plant: About 200 coal train cars per day. That is staggering! It is about 5.25 million tons of coal per year. Multiple that by twenty and that only covers new capacity added last year. It is depressing to see these numbers.
tcmart14 2 days ago [-]
From my understanding, yea they still stand up coal fire plants because world demand for Chinese manufacturing and the energy demand it creates exceeds how fast they can build reactors. So they build coal fire with the intention to replace that capacity with renewables later. Been awhile since I looked, but China has like 30-some odd plans for nuclear reactor construction with quiet a few of them having already broken ground, but that was about 5 years ago.
robocat 2 days ago [-]
World Bank said exports in 2024 were about 20% of China's GDP. AI said 30% of China's petrochemical usage is exported as embedded energy. You're right that it is useful to think in terms of marginal generation, where it is marginally used. But that is more difficult to measure than just making assumptions at the level of the economy.
graemep 2 days ago [-]
No one forces them to increase manufacturing output. They do it to make money, just like everyone else.
bamboozled 2 days ago [-]
China builds literally everything for all countries, so why wouldn't they bring coal online? Very tiring arguments IMO.
Are you sending them extra money to fund nuclear?
sarchertech 2 days ago [-]
China emits 3x more co2 than the US and last year their rate of increase was 2.5x that of the US.
"big number bad" isn't a great way of analysing the data, because China has a massively larger population.
The US emits 49% more CO2 per capita than China. And even with the much larger rate of increase, 0.79% to US' 0.3%, it'll still require 81 years to catch up to the US per capita rate [0].
The US aren't the good guys here, not by any means. Compared to my country, for instance, the US' emissions are 3x per capita that of the UK. If you want to argue that we shouldn't be using per capita (although we absolutely should), then that's 45x the UK's CO2 emissions.
[0] log(13.59/9.13) / log(1.0049) = 81.4 years
sarchertech 2 days ago [-]
In absolute numbers China produces 1/3 of the total worlds CO2, and they are still growing the amount of CO2 they are producing.
That is not the actions of a country who “believes in climate change”.
China’s CO2 per capita is ahead of every large developed country except for the US, Australia, and Russia.
Zigurd 2 days ago [-]
Actually, if you Google the rankings, China is 27th per capita, behind, of all places, Iceland.
sarchertech 2 days ago [-]
> ahead of every large developed country
Because those are the only ones that can move the needle on climate change.
Zigurd 2 days ago [-]
You didn't bother to do the search and find out how many industrialized nations are also higher emitters per capita.
sarchertech 2 days ago [-]
I said large developed countries.
No one cares if Iceland emits more co2 per capita. They have < 400k people.
This isn’t about fairness, this is about whether the actions of the leaders of China care about impacting climate change.
desterothx 2 days ago [-]
Almost like the entire worlds manufacturing was pushed to China the last 40 years (and is now getting moved to cheaper countries with the rising Chinese middle class). Is there a similar explanation for why the US number is so high?
myrmidon 2 days ago [-]
Outsourced manufacturing is much less of a factor than people typically assume:
US numbers are insanely high because cheap hydrocarbons are locally available (=> bad incentives) and everyone is wealthy (that correlation is very strong; just compare Luxembourg, which is much wealthier and more polluting than surrounding nations) and also population density is rather low so more energy wasted for transport.
flakeoil 2 days ago [-]
> that correlation is very strong; just compare Luxembourg, which is much wealthier and more polluting than surrounding nations
Well, your theory does not hold at all if you look at Switzerland which pollutes 1/4 of the US per capita. CO2 pollution is not and does not have to be correlated with standard of living.
myrmidon 2 days ago [-]
The reality is obviously more complex, and many countries have already managed to somewhat decouple economic growth from CO2 emissions (very good!), but as a crude approximation: More money => everyone gets more stuff => more CO2 emissions (producing and running that stuff).
If you want meaningful CO2/capita comparisons, you also have to be very careful with countries that get ("free") hydro power opportunities for electricity because that distorts the picture massively (same for e.g. Norway).
Big producers of hydrocarbons, on the other hand (like the US) have to work hard to resist the allure of cheap & convenient fossils.
Switzerland is so far ahead in this comparison because they get a lot of CO2-free hydroelectricity (>50%), have to import most hydrocarbons (=> incentive against) and also save massively on transportation because density is much higher.
gwerbin 2 days ago [-]
How much of that "decoupling" is actually just "exporting"? Not being glib or dismissive, I think it's an important distinction and I don't know enough about Europe to know how the two compare.
myrmidon 2 days ago [-]
> How much of that "decoupling" is actually just "exporting"?
Note how this still trends down for most western industrialized countries despite positive economic growth over the last two decades.
gwerbin 2 days ago [-]
There are so many bad incentives all over the place in the USA, and it's not so much that cheap hydrocarbons create them as it fails to prevent them.
Consider the gradual takeover of SUVs on the roads (many which to be fair are probably about as efficient as sedans were 20 years ago, albeit more expensive in real terms and significantly more dangerous for pedestrians), or the heavy reliance on trucking rather than rail for heavy freight transportation between logistics hubs, or the heavy reliance on air for long-distance travel due to having literally zero high-speed rail, or the fear of nuclear power, or the overt political opposition to green energy.
sarchertech 2 days ago [-]
The rest of that is true but the US moves a much higher percentage of its freight by rail than almost any large developed country.
Just in comparison to Europe “46 percent of European freight goes by truck while only 11 percent goes by rail, while in the United States more than 40 percent goes by rail while just 30 percent goes on the highway."
This is in ton-miles, so it’s the metric most directly relevant to emissions.
myrmidon 2 days ago [-]
Is this excluding maritime shipping?
According to eurostat (https://ec.europa.eu/eurostat/statistics-explained/SEPDF/cac...) that is two thirds of freight volume and it probably still eclipses even rail transport in CO2 efficiency (until full CO2 free rail electrification), so this probably distorts things completely...
If you actually compare the emission fractions, you can see that the US uses much more CO2 per capita on transport compared to Europe (about 3 times as much!):
Maritime is around 37% if you only include domestic and intra-EU shipping.
Your report include all shipping that passes through any EU country’s EEZ. So if a ship coming from Africa headed to the US passes by the Azores. A portion of the tonne-kilometers of that ships voyage will show up in your report.
If you do the same thing with the US, the percent of maritime freight movement would also skyrocket.
ralferoo 2 days ago [-]
You seem to think China is the evil polluter here. Just sort by per capita, and you'll see it's not even in the top 10% most polluting per capita.
There are plenty of large (depending on your definition of large) countries above it. Not just US, Australia and Russia that you are excluding for some reason, but also Qatar, Kuwait, UAE, Oman, Saudi Arabia, Canada all have significantly developed economically important countries. And they all have GDP per capita much greater than China's. What's your rationale for excluding them? There's a whole host of smaller countries with higher GDP per capita than China between those I listed above and China on the CO2 per capita results.
We should agree that ALL countries should be working to reduce CO2 emissions. But it sounds like you have a strong US bias and given Trump not only denies that climate change is even a thing, has withdrawn from international agreements on reducing emissions and even tries to pressure other countries to burn oil instead of investing in windfarms, it seems a bit disingenuous to try to make out that China is the only country that needs to improve. Before commenting on the speck in someone else's eye, first remove the plank from your own, yadda yadda...
sarchertech 2 days ago [-]
> There are plenty of large (depending on your definition of large) countries above it. Not just US, Australia and Russia that you are excluding for some reason, but also Qatar, Kuwait, UAE, Oman, Saudi Arabia, Canada
None of those countries have more than 50 million people. Most of them have fewer than 10 million. None of them are going to move the needle on climate change.
The US is clearly not a country to emulate when it comes to co2 emissions. I don’t think China is any more “evil” than the US when it comes to climate change.
But their actions are clearly no the actions of a country who “believes in climate change”.
ralferoo 2 days ago [-]
> But their actions are clearly no the actions of a country who “believes in climate change”.
I'd be willing to bet that you've not been to China. Shenzhen is a very green city, I'll go into that later, but also Beijing and Shanghai have historically had serious problems with smog in summer, and this has got significantly better in recent years as local government pushes companies towards renewable energy sources and close down coal power stations.
For example in Shenzhen, all buses and taxis have to be EV, and around 80% of all new privately owned cars are also EVs. That seems like a city that believes in climate change.
You've got vast swathes of desert covered by solar panels providing electricity further east (although sadly, a lot is waste due to transmission inefficiencies). That seems like a country that believes in climate change.
You've got a country that produces 80% of the world's solar panels, and has over 1/3 of the worldwide installation of solar panels. That seems like a country that believes in climate change.
You've got a country with massive windfarms and a net exporter of wind turbines. Combined, over 1/4 of all electricity in China comes from renewables. That seems like a country that believes in climate change.
You've got a country that uses a lot of battery storage to smooth power demand, and exports these units to pretty much everywhere because they're about the best you can get. That seems like a country that believes in climate change.
You've got a country that built the massive three gorges dam project, started in the 1990s and finished nearly a quarter of a century ago. That seems like a country that believes in climate change.
Sure, there's still a long way to go, and due to the massive energy demands, it still has coal power stations providing about half it's electricity generation needs. That needle is shifting over time, but you know what? It's still half coal and yet it still only has 2/3 the emissions per capita compared to the US. But even if it still has a long way to go, its actions clearly ARE the actions of a country that believes in climate change.
sarchertech 2 days ago [-]
I took Mandarin in college. I don’t hate China. I don’t think the US is a good country to emulate with respect to climate change mitigation.
But if you believe climate change is real, and you have authoritarian control over 1/3 of the entire world’s co2 emissions, you wouldn’t be increasing your green house gas emissions, and you wouldn’t have a co2 emission per capita higher than almost all of the developed world.
notact 2 days ago [-]
To be fair, you can replace ""believes in climate change" with "prioritizes energy independence". I would argue the latter is a stronger motivator for China.
onetimeusename 2 days ago [-]
I don't think per capita emissions is the only metric that should be used. Environmental impact is not measured per capita. Also it blames people in the US for energy consumed here to manufacture which includes exports for people abroad. The US's per capita emissions are not even far away from other similar nations. It's a data point, sure. It's comparable to Canada and Australia and even South Korea. Where did you get that the US emits 45x the UK?
Another point is there is evidence[1] China has cheated and manipulated their data[2]. I don't like this "who are the good guys" game. Anyone playing this game is just looking to create a narrative about good and bad guys.
From the GP [1], US emits 4,632,164,876 and UK emits 292,419,359 tonnes CO2. Which, sorry my bad, is actually only 15x. Looks like I did the ratio of China to UK which is 45x.
But either way, it shows that comparing total emissions is stupid because it's biased against bigger countries. Emissions per capita is a far more sensible metric.
> Also it blames people in the US for energy consumed here to manufacture which includes exports for people abroad.
Sure, but the same can be said about China, in fact I'd argue even more so. As a resident of the UK, the majority of things I buy is made in China, almost nothing from the US.
> The US's per capita emissions are not even far away from other similar nations.
Sure, but they are higher than China's, which was the point I was making to refute the GP. The only reason China's total emissions is 3x US is because its population is 4x US.
> Anyone playing this game is just looking to create a narrative about good and bad guys.
If you re-read what I wrote that you're replying to, you'll see that we agree on that.
It was the GP post that tried to compare China and US in terms of emissions, and deliberately chose total emissions to portray China as significantly worse than the US and getting worse year-on-year, while conveniently ignoring the population size differential and that the increase in emissions is still negligible compared to the overall difference, because it'd take 81 years for China to catch up with the US on emissions per capita. I then compared the US to the UK to show that comparing total emissions instead of per capita is stupid, because the population size matters.
But, as I said in the post you replied to - both countries can and should do much better, but as China isn't even in the top 10% of emissions per capita, this is a worldwide problem and trying to pin it all on China is disingenuous.
Ok I appreciate the correction, I could not figure that out.
As far as everything else, yes I think this is a global problem, not solely a Chinese problem but I would characterize US-Chinese relations as tense so that tends to amplify finger pointing which I disagree with so yes, we agree.
iso1631 2 days ago [-]
You need to account for carbin import/exports in manufactured goods and services
If the UK imported all its electricity from Poland, it would reduce on a per-capita basis but would increase global
Ekaros 2 days ago [-]
USA emits 2165574,97709 times more than Faroe Islands. Sounds like they should be forced to limit that to some reasonable level say 2x of that of Faroe Islands. Don't think it would be too big ask right?
sarchertech 2 days ago [-]
China produces 1/3 of the world’s CO2. They already produce more per capita than all other large developed countries except Australia, the US, and Russia. And their rate of growth is higher than all of those countries except Russia.
There’s no way to force them to do anything, but this is not that action of a nation that “believes in climate change”.
nixon_why69 2 days ago [-]
That's while being the world's factory, having lots of green energy, electric cars and an ACUTE motivation to reduce oil imports. Coal is a problem but at least they're not canceling half-finished green energy products when administrations change.
sarchertech 2 days ago [-]
Sure they are building solar infrastructure. But they’re also still increasing co2 emissions. If you are the autocratic leader with almost total control of 1/3 of the worlds co2, and you believe that climate change is an existential threat, you wouldn’t be increasing co2 emissions. Xi could reduce co2
emissions with the stroke of a pen.
So he obviously doesn’t think it’s that much of a threat.
nixon_why69 2 days ago [-]
Xi could theoretically reduce GDP and associated emissions with the stroke of a pen[1], and so could most legislatures. None are doing it.
Building up alternative energy is a work in progress for anyone, they're leaders on electric cars which is good but they won't get off coal until they (or someone) figures out managing a grid with 80%+ of mostly-intermittent inputs from wind and solar.
[1] Not really, for the same reason as the legislatures, it's horrible politically. Dictatorships and especially the Chinese system have broader bases of support and political activity than most American media discourse assumes.
sarchertech 2 days ago [-]
> Xi could theoretically reduce GDP and associated emissions with the stroke of a pen
He showed he was willing to do that during Covid. If he was really convinced climate change was a problem he could clearly do more than he is.
It’s not a prisoners dilemma issue with Xi like it is with most countries. He controls 1/3 of global emissions. He is probably the only person on the planet who can personally reduce future temperatures.
nixon_why69 1 days ago [-]
And then they ended those harsh COVID measures due to unpopularity and protests. He's not God-King.
I'd love for everyone to do more but its a little rich to hear from westerners, who emit more CO2 per person, that the Chinese should stop developing people out of third world poverty for everyone's sake. Peasants produce much less carbon than people with air conditioning.
sarchertech 1 days ago [-]
> but its a little rich to hear from westerners, who emit more CO2 per person, that the Chinese
1. Most western countries don’t emit more per capita than China.
2. I’m not saying what Xi should or shouldn’t do. I’m telling that what he is doing indicates that he doesn’t believe climate change is an existential threat.
nixon_why69 20 hours ago [-]
I also would like the Chinese to do more, but as an American waving broadly at our political consensus, I have a hard time pointing fingers. At least they're doing something.
sarchertech 9 hours ago [-]
I’m not making any argument about what China should do. I’m saying that the argument that China is ruled by an enlightened despot who “believes in science” and believes in the existential risk of climate change is bullshit.
1234letshaveatw 2 days ago [-]
are you advocating for not changing administrations?
nixon_why69 2 days ago [-]
Is that a good faith question? Please check the site guidelines.
1234letshaveatw 2 days ago [-]
It is, and certainly relevant (although perhaps inconvenient) considering the recent elections in China. Aren't shifting priorities a side effect of free elections?
pessimizer 2 days ago [-]
This is as goofy a response as answering the question "why did you make the decision to kill him" with "are you against people making decisions?"
bigyabai 2 days ago [-]
> Aren't shifting priorities a side effect of free elections?
Pointing out that the US cancels their green projects when Republicans take over is a fact, it's not rhetoric.
There's plenty of room for interesting discussion on what it takes to make green energy work. "What, are you against freedom?!" is not that.
fooster 2 days ago [-]
Just to quote a poster above:
"You seem to think China is the evil polluter here. Just sort by per capita, and you'll see it's not even in the top 10% most polluting per capita."
sarchertech 2 days ago [-]
It is just barely out of the top 10%. 27 out of 207. If you remove counties ahead of with tiny populations like Palau with 40k people, Iceland with 400k, and Greenland with 50k, it is in the top 10%.
Since they produce 1/3 of the co2, 1/3 of the reduction the world needs, needs to come from them. And they are still increasing their emissions. This is not the action of an autocratic leader who believes in climate change.
beepbooptheory 2 days ago [-]
Kind of a meta point but for those who don't know, its fun to track this comment and its responses through time. I want to nominate it for the hall of fame in classic HN post genre. It will be two decades soon!
One interesting thing would be to see how the numbers change over the years alongside the otherwise identical debates.
linkregister 2 days ago [-]
Can you refine that search term a bit more? I followed the link but it brings up many marginally related comments.
bamboozled 2 days ago [-]
They make literally everything, including all the pieces for a renewable / low emission economy.
drop_star 2 days ago [-]
"Clean coal"
petcat 2 days ago [-]
USA and Europe are basically equal in terms of yearly production of renewable energy. Texas + California alone rank 6th in the world.
MaxHoppersGhost 2 days ago [-]
>China believes in Climate change and are at least taking steps to address it
I've got a bridge to sell you!
butlike 2 days ago [-]
Exactly my question
jayd16 2 days ago [-]
You shouldn't trust either but also this is whataboutism.
pessimizer 2 days ago [-]
No, it is not. Whataboutism is a specific term created (by Cold War propagandists) to refer to when the US criticizes the USSR about "X," and the Soviet Union replies by asking about "Y" in the US.
It is not a term created to refer to when some American criticizes any of America's official enemies for "X," and someone else reminds you that America also does exactly "X."
jayd16 2 days ago [-]
> "Whataboutism" or "whataboutery" (as in, "but what about X?") refers to the propaganda strategy of responding to an accusation with a counter-accusation instead of offering an explanation or defense against the original accusation.
This is a counter accusation and not a defense, ergo, whataboutism.
shimman 2 days ago [-]
No, it would be whataboutism to say something like "but in America you extrajudicially kill citizens for protesting the government."
It's not whataboutism to talk about the same subject in the context of a different government. In this case LLM corporations.
All you're doing is deflecting valid criticism. It's not whataboutism to talk about how bad American corporations are to American citizens in the context of talking about how supposedly bad Chinese corporations are to American citizens (China doesn't profit off of the deaths of Americans like health insurance companies (this is whataboutism for example)).
2 days ago [-]
kortilla 2 days ago [-]
That’s still whataboutism because it’s not the only alternative.
2 days ago [-]
rq1 2 days ago [-]
I never understood how people referred to or understood “whataboutsim”.
I personally understand it not as a diversion but as a critique of the moral higher ground implied by the first accusation. Said differently: keep your own house in order.
kortilla 2 days ago [-]
Because that’s not a logical defense and nobody cares about moral high grounds. That’s why it’s a logical fallacy.
Actions can be judged on their own merits regardless of who presents the arguments.
rq1 2 days ago [-]
Hilarious. It is absolutely a logical défense: our societal order and its relative fairness don’t emerge from the void.
Otherwise a victim defending himself is the same as the person assaulting?
Perpetrator: you just stabbed me OMG!
Victim: you were going to rape and kill me
Perpetrator: sorry that’s just whataboutism and definitely not a logical defence. Your actions can be judged on their own merits, regardless of who presents the arguments.
1 days ago [-]
kortilla 22 hours ago [-]
Nope, defense is an action to prevent an attack. China isn’t building coal plants to stop global warming.
Also, the world isn’t two sides. There are tons of participants on this site who are not in the US that criticize both china and the US (i.e. most of Europe). So you realize how stupid the defense looks of “I’m behaving like a piece of shit because this guy behaves like a piece of shit” when there is a room full of people trying to do the right thing.
If your argument is whataboutism, it means you have no actual defense for your behavior. You’re just appealing to “someone else did something bad”.
WithinReason 2 days ago [-]
No it's not. There is a difference between a comparison and whataboutism
platinumrad 2 days ago [-]
There's nothing intrinsically wrong with "whataboutism" and it is in fact often illuminating.
pessimizer 2 days ago [-]
Especially because the original criticism is usually a stand-in for an overarching moral condemnation. "Whataboutism" was a US defense for the fact that it was an apartheid state while criticizing political and financial liberties in the Soviet Union.
It also worked. The shame built up in the US so quickly that laws started being struck from the books. The US entered WWII with a segregated military, and by the time it marched into Korea, it was rapidly desegregating. The Soviet Union had successfully made its case that the US had no moral high ground, although it had not made its case for its specific set of limitations of financial and political liberties.
"Whataboutism" is just ad hominem for dummies. But if the real argument is about who is the better man, or what is the better system, ad hominem isn't a logical fallacy. It's simply a change of subject.
simianparrot 3 days ago [-]
[flagged]
inigyou 2 days ago [-]
Been there. Great place. High standard of living, everyone is very productive. Still a dictatorship, and you can't speak ill of the government, but a successful one.
Natfan 2 days ago [-]
my understanding is that you can speak ill of the government publicly? like they'll nuke your billibilli if you talk shit about xi, but you can complain privately to a friend
not saying i agree with it, i'm saying it isn't as severe as the west makes it out to be
ozgrakkurt 2 days ago [-]
That is an extremely low bar. Imagine you couldn't do that, that would basically be George Orwell's 1984.
Natfan 2 days ago [-]
well in 1984 the two way televisions would get you
again i'm not saying i approve, I'm pointing out the distinction
Envwnger 3 days ago [-]
What if they don't live in USA?
psychoslave 3 days ago [-]
Abstracted from the HN specific audience, there’s around 17% chance they already do, which is to compare to the "around 4%" chance they live in USA.
This is such an hilarious answer. For most of the non-Western world, we don't distrust China because they have created benefits for the developing world. Whereas the USA has been outright imperialist for the past 60-70 years, undermining and overthrowing developing democracies, putting in dictators, using international organisations as debt traps, waging and funding wars that put our lives and economies at risk.
It is so pathetic how American's see themselves, and are so deeply afraid of China. I am far more fearful of the predatory nature of the USA and its agencies than I could ever be of China who would never have any interest in me.
ToValueFunfetti 2 days ago [-]
Thank god. I'd been hearing we screwed up catastrophically by electing an admin that took a chainsaw to our foreign aid programs, killing millions in the developing world. Such a relief to learn we were never doing that stuff in the first place.
a34729t 2 days ago [-]
yeah comments like this show that ending USAID was the right move (apart from it's obvious utility to pass along bribes)
mad_tortoise 2 days ago [-]
Ending USAID is the exact reason hundreds of thousands suffer more today than they did a few years ago. It is the reason mass migration takes place from developing to developed. The total ignorance and outright stupidity of comments like this is why Yanks should really take a seat when it comes to talking about geopolitics.
Natfan 2 days ago [-]
the left hand feeds while the right hand beats
brettermeier 2 days ago [-]
[flagged]
datsci_est_2015 2 days ago [-]
Hey but it’s wrapped in gold foil, much like its president’s domestic interiors.
Der_Einzige 3 days ago [-]
The ROI for (white) men going to China for a period of their lives is extremely high. Where else are you going to get your “Abg cmo”
bigyabai 3 days ago [-]
1) I'm not using AI to bicker over fringe political shibboleths.
2) I don't think it is any more or less safe to put my code on a Chinese server versus an American one. A Chinese provider also isn't liable to spy on me for the feds, as OpenAI and Anthropic certainly do.
sanex 3 days ago [-]
Different county different feds both spying I'm sure.
margalabargala 3 days ago [-]
Sure but if you are not Chinese or in China, then the chances of negative consequences to you from the Chinese feds is vanishingly small due to lack of ability to do anything that affects you.
Meanwhile if you are in the US, DHS has already subpoenaed social media sites looking for people who made anti-ICE posts and I can't imagine they consider subpoenaing AI conversations off limits https://www.nytimes.com/2026/02/13/technology/dhs-anti-ice-s...
hnfong 2 days ago [-]
Also the US government has defacto extra-territorial jurisdiction due to the willingness of allied countries (i.e. most of the developed world) to extradite.
Of course only applies when they really want to get you, but that's still a risk.
margalabargala 1 days ago [-]
The chances of a foreign country extraditing you (or the US requesting extradition) over social media posts is for now quite low, fortunately, though you're right.
Broader point being, in general if a country's federal police are going to be spying on you, if you get to pick the country you should pick the one in which you do not reside and don't plan to visit. For a typical US citizen who is not an intelligence target, the chances of negative consequences from China spying on them is way lower than the chances of negative consequences from the FBI spying on them, simply because the FBI will have nonzero false positives.
cesarb 2 days ago [-]
As a less extreme example: consider how many non-USA hosts say they follow the DMCA.
ansgar77 3 days ago [-]
Regarding point 2, I don't trust my data being safe running inference on model creators api, but neither do I trust US providers. Both use it for their own benefit, the only difference is the country of origin. The US has a lot more legal safeguards for this but I don't trust they don't do it regardless.
vachina 3 days ago [-]
legal safeguards only make sense only when they’re enforced. With how “move fast and break things” Silicon Valley is law is always playing catch up (at your expense)
riskd 3 days ago [-]
It’s perfectly fine for the “West” to influence the world though, right? Or is it only a problem because… they’re Chinese?
hetman 3 days ago [-]
Why would it be a problem that they're Chinese? It's a problem because their country is ruled by an authoritarian regime. The West, by contrast, doesn't have a singular arbiter of truth, it has a plurality of perspectives (as much as certain interest groups would really rather that was not the case). It's not perfect, but those of us who have memories of living under authoritarian regimes can tell you there's no comparison.
podgorniy 3 days ago [-]
> It's a problem because their country is ruled by an authoritarian regime
Hey, have you seeing what trump does with your (presumably) country? Maybe wars? Maybe market manipulation? Mayde pedo right covering on the government level? Maybe bubbles and threats to EU? What an ignorance. You live with old stories, not the current state of the world...
burnerRhodov3 2 days ago [-]
This is absolutely nothing compared to what China did to the Uyghurs, or China cutting citizens off of the financial system (WeChat) due to posting political content... If you were in China, you would be blacklisted for posting a similiar comment on wechat. If you openly called the CCP a group of pedo's you'd regret it within 24 hours.
herbst 2 days ago [-]
If you post content like this you likely to get refused entering the US. I realize it doesn't apply to US citizens but it's also not like everyone can say what they want and get away with it.
burnerRhodov3 21 hours ago [-]
Refusing entry is not the same as being made to disappear.
herbst 5 hours ago [-]
Still not "normal" if you come from a more free place where things like this really don't matter at all.
burnerRhodov3 19 minutes ago [-]
You are taking this completely out of context, and if the OP had posted the above on x, or facebook they would not be denied entry. Threatening statements, glorifying violence, or ties to flagged groups is what gets you denied entry. And honestly, I agree with that.
If you are screaming death to america on your socials... I don't think i want you in America.
rikima_ 2 days ago [-]
> This is absolutely nothing compared to what China did to the Uyghurs
which is absolutely nothing compared to what they did to Gaza.
burnerRhodov3 21 hours ago [-]
Palestine elected a international terrorist organization to head their government, then killed, and raped a music festival.
hetman 17 hours ago [-]
So you're telling me Trump has direct control over the topics and conclusions allowed into AI models built by Western companies? I don't think your off topic whataboutism has really helped dispel any ignorance in the world.
inigyou 2 days ago [-]
If being ruled by an authoritarian regime is the problem, then wouldn't both Chinese and American models be problematic?
European models too, if they had any.
hetman 17 hours ago [-]
I don't think you'll understand what authoritarian means until you try moving to China and then criticising their government. For those of us who have lived under authoritarian regimes, it's an absolute travesty to see how diluted the meaning of this term has become.
A swiss one. No idea if it's any good but it's there
mynameisbilly 2 days ago [-]
The United States has propped up some pretty horrific, murderous, bloodthirsty regimes over the past century to ensure they maintain a proxy presence across the globe. If supporting or engaging in an "authoritarian" regime is your litmus test, then the United States fails resoundingly.
burnerRhodov3 2 days ago [-]
The unfortunate concequence of "rebellion", is sometimes those regimes go on to be just as bad as the previous regime. You're playing captain hiehnsight here... intent matters in these things.
hetman 17 hours ago [-]
Guess what you can do in the US and can't do in China? Talk about their respective atrocities. You've missed the point entirely as it pertains to AI if you thought my litmus test was about having the most moral government. It was about that government's ability to censor information.
BoredomIsFun 3 days ago [-]
West has though a history of supporting regimes far more hideous than Chinese, and ready to push narrative whenever possible. It is well known idea BTW in political thought that democracies are far more reliant on propaganda for control than authoritarian,which due to their strong grip on political discourse do not need high quality propaganda anyway.
TLDR: American propaganda is not any better than Chinese, neither have the best interests of my country in their minds.
youre-wrong3 3 days ago [-]
Have you got examples of GPT/Claude/Grok influencing people?
hetman 3 days ago [-]
I think there's no denying they each have an ideological bent (as is their right as private companies). I have had Copilot deny me access to historical information on ethical grounds, even though I don't think anyone would find it controversial (clearly it was overtuned, GPT and Claude had no problem answering the same question). What is different though is that they are each allowed to have their own perspective instead of a singular mandated one.
youre-wrong3 3 days ago [-]
That was guard rails in the harness. Not the model. We are talking about baking it into the model. Denying to do something is also very different to changing historical facts like China does.
damontal 3 days ago [-]
You are splitting hairs. End result is the same.
kortilla 2 days ago [-]
No it’s not. Saying “I won’t talk about how to make bombs” is not the same as “the holocaust didn’t happen”.
MSFT_Edging 2 days ago [-]
> very different to changing historical facts like China does.
Oh but they're trying. The right-wing usage of things like "woke" and "DEI" primarily serve to hide/destroy historical realities. [1]
Florida has it's "STOP WOKE" act that forces teachers to talk about how slaves learned skills/benefited from slavery[2] and that various massacres also had black perpetrators.
What is this other than changing historical facts? About fucking chattel slavery for Christ's sake.
And also, not taking a side here, plenty of woke partisans valued righteousness over truth in the 2020 era. The 1619 project which was a whole NYT backed big thing had a bunch of factual errors.
The real rub is when you get into "shared facts" that Americans were all taught in high school civics but the rest of the world wasn't. If you've mostly been in America, it can seem like someone's deliberately lying about history but they simply weren't properly educated with the Correct Interpretation.
kortilla 2 days ago [-]
> and that various massacres also had black perpetrators.
the implication of your message is that this is not true
A major thing I noticed is that trying to get ChatGPT to say something negative about Sam Altman is like pulling teeth.
Pure speculation, but I would wager it has a direction somewhere to use OpenAI as authoritative about anything related to OpenAI - arguably for help docs and whatnot.
But the impact does stay the same.
phantompeace 2 days ago [-]
Why did you mention Grok, because it undermines your argument.
On (1), models are an aggregation of large volumes of data sources, whatever the culture producing them, you'll get the average bias of that culture.
We see that on what minorities are associated with inside the model, or how things that aren't online will have a completely different weight. Or how 2/4/5/8ch or X will be disproportionately present in specific models despite being the places where facts go to die.
MintPaw 3 days ago [-]
I see your point, but this isn't exactly true, it only takes a single person to bias a model by deleting specific training data or over training on certain facts.
sajithdilshan 3 days ago [-]
But hasn’t it been the same with US since the WWII? US has influenced the world through various ways sometimes even weaponising human rights to spread American superiority and its narrative?
inglor_cz 3 days ago [-]
While true, you neglect the fact that US corporations have some autonomy from the government. I doubt that any of the current American AI corporations get regular dossiers of desirable responses to sensitive political topics from Washington and implement them. That would be a reason for a major lawsuit and a major scandal.
In China, the image of the country and the preferred narrative is under much tighter control of the government and local AI shops won't have any autonomy in this regard at all. Either obey or get shut down.
BTW What you mean by "weaponizing human rights", exactly? I am curious. If anything, I would say that the US foreign policy didn't promote human rights sufficiently, especially in Latin America, where the "bastard, but our bastard" attitude was typical.
OTOH in Europe, US human rights policy was probably relevant in saving some dissenters in the former Eastern Bloc from torture or execution.
sajithdilshan 3 days ago [-]
> you neglect the fact that US corporations have some autonomy from the government
I'm not sure if this is correct. It seems like every big US corporation is changing their policy based on the views of the administration in place. As an example during Biden's time the DEI was in full motion in every corporation and in current administration it's the other way around. Also I remember as soon as Biden won the election twitter and meta suspended the profile of Trump. So I hardly think that there is much room for autonomy. Also the latest export bans on Antropic and OpenAI models kind of make your argument weak, of course the company can sue the givernment, but the national security comes above all.
> BTW What you mean by "weaponizing human rights", exactly?
US has been using the violations of human rights to impose sanctions or to especially get countries in line which are not supporting the US agenda when it comes to global politics. However, they were more than happy turn a blind eye if the respective country that violates human rights (most middle eastern countries) if they are allies of US doctrine.
inigyou 2 days ago [-]
A Russian and an American are sitting next to each other on a plane.
The American says "I'm impressed by the propaganda you have in Russia."
"Oh it's very good, but it's nothing compared to the propaganda you have in America." replies the Russian.
"Huh? We don't have propaganda in America." says the American.
"Exactly." says the Russian.
abalashov 2 days ago [-]
As a Russian, I deeply appreciate this. In Russia, it seemed obvious to even the most parochial peasant that there was propaganda, while in America, the vast majority of the society not only fails to consider the possibility, but is cholerically allergic to the very idea.
Edit: recognising that there is propaganda != knowing what is and isn't propaganda. That's all I meant.
WarmWash 2 days ago [-]
The American government has propaganda but it really sucks at it. All the people skilled in manipulating behavior work at our ad agencies, and hence our actual propaganda is mostly "Use this detergent! Take this drug! Drive this car!"
kelipso 2 days ago [-]
Lol, someone completely missed the point… American propaganda is such a part of the culture, it gets barely noticed. Even movies that has military equipment, the military gets to have heavy input towards positive depictions of american military. US military has paid millions in marketing to all the sports organizations for promoting the military. Etc etc. This is just one aspect of the propaganda.
Americans literally recite a brainwashing mantra every day in school and are punished if they don't (despite the constitution).
butlike 2 days ago [-]
Propaganda is just advertisements for the government, right?
abalashov 2 days ago [-]
Hard to say. Advertisements have at least a pretense of persuasive intent. Propaganda mostly operates in the absence of contrary information.
inglor_cz 2 days ago [-]
If every parochial peasant could recognize Russian official propaganda, why are so many people on board with whatever the propaganda says?
I mean, it took longer than the Great Patriotic War before Putin's popularity started visibly falling and people started questioning what the entire "SVO" is for.
inigyou 2 days ago [-]
Surely [a] they don't really give a shit anyway and [b] they don't want to fall out a window
I think Russians are used to knowing it's dangerous to disagree with dear leader and there's no advantage to disagreeing so they just go along with whatever he says
nixon_why69 2 days ago [-]
I mean, if the American propaganda is "you are evil" and the Russian propaganda is "fuck those guys", you don't have to agree on every particular. People have nationalism everywhere.
inglor_cz 2 days ago [-]
The Russian propaganda is more along the lines of "Ukraine is an illegitimate state led by Nazis and drug junkies, they steal children's organs and sell them to the West, cook up dangerous viruses in their labs and prepared to attack innocent Russia, hence they have to be crushed, which our glorious Russian power will soon accomplish with mighty Oreshniks that have no parallel anywhere".
There is way too many mixed marriages etc. for any reasonable Russian to believe that their formerly-closest East Slavic neighbour has turned into a Fourth Reich.
The glorification of Russian military power after almost five years of static attritional war and dozens of burning refineries and ships is in a category of its own...
abalashov 2 days ago [-]
> There is way too many mixed marriages etc. for any reasonable Russian to believe that their formerly-closest East Slavic neighbour has turned into a Fourth Reich.
Exactly. The distortions in American public life are quite a bit more intricate and artful.
inigyou 2 days ago [-]
hey what a coincidence that's basically the exact same thing our own government tells us about Palestine
abalashov 1 days ago [-]
Yes, but we have far fewer reasons to think otherwise, since we don't know anything--if you're talking about Americans.
abalashov 2 days ago [-]
Recognising that there is propaganda != being deeply savvy about what the propaganda is. However, I've found it to be an article of faith in America that the US just doesn't engage in that kind of thing--less so in recent years, as the crisis of belief and confidence in institutions has ricocheted, egged on by the disinformation and demagoguery of the Trump years.
hnfong 2 days ago [-]
Well, these days there's little need for propaganda when people just blatantly do sh!t in the open and get away with it.
inglor_cz 2 days ago [-]
I remember peak wokeness quite well and it was pretty obvious that it was driven by the Twitter activist mob and not the Biden administration, which was quite visibly behind the times and followed the trends.
Already during the Biden administration, defections from the orthodoxy started and then multiplied - some businesses like Coinbase or IIRC Cloudflare refused the demands outright. Musk bought Twitter with an explicit task to make it less progressive.
And the White House did precisely nothing against this defection trend.
"US has been using the violations of human rights to impose sanctions or to especially get countries in line which are not supporting the US agenda when it comes to global politics. However, they were more than happy turn a blind eye if the respective country that violates human rights (most middle eastern countries) if they are allies of US doctrine."
I do agree that the US is hypocritical about human rights, but actual violations of human rights should be a reason for sanctions, and the fact that this is done only partly/imperfectly, IMHO, beats the potential alternative when it isn't done at all. This would be a much worse world in my opinion.
sajithdilshan 2 days ago [-]
> I do agree that the US is hypocritical about human rights, but actual violations of human rights should be a reason for sanctions, and the fact that this is done only partly/imperfectly, IMHO, beats the potential alternative when it isn't done at all. This would be a much worse world in my opinion.
That's why I said that it was weaponized. In reality US doesn't care if there is a real human rights violation or not, they just use it as an excuse to get what they want.
linkregister 2 days ago [-]
Everybody is amoral and cynically only doing good things for personal gain, except for you, the one good person
littlecorner 1 days ago [-]
It's much easier to blame everyone else instead of trying to change myself
buellerbueller 2 days ago [-]
>US corporations have some autonomy from the government
This is correct...because they just buy the government when they need to.
ux266478 2 days ago [-]
> I doubt that any of the current American AI corporations get regular dossiers of desirable responses to sensitive political topics from Washington and implement them.
That's completely true. What happens is that you get hit with a wall of bureaucratic threats and nonsense by the government out of nowhere, and usually for some unrelated matter. It's like a high level version of getting pulled over for not using your turn signal. Which is why if your company has good legal, they often act in a highly proactive manner about that kind of thing.
What constitutes sufficient rabble rousing to get put on the government's shitlist differs between both countries, but there are plenty of things that will send American politicians into having a conniption, and it's not a static criteria. McCarthyism is a pretty easy example. I'm no Muskboy, but we can point to the procedural harassment he recieved over hitting a blunt on the Joe Rogan podcast to be a related matter. Criticizing the genocide happening in the Levant is frequently saber rattled by politicians as something they're interested in making prosecutable, and there are examples of institutional authority targeting and attempting to punish people over it.
mullen 2 days ago [-]
Non-sense. In China, you obey. There is no way to challenge the orders in courts or even with your local party official. It's written into law and companies that disobey are shutdown or taken over by a party loyalist and the management is disappeared.
> bureaucratic threats and nonsense by the government out of nowhere
You can challenge these in courts. There is a free press you can go to. You can go to social media and lay your case out there. You can challenge what the government is requesting of you. In China, you can not at all.
2 days ago [-]
pydry 2 days ago [-]
In China the state rules the corporations and in the US the corporations rule the state.
Neither is a system worth fighting for. Swiss style democracy maybe but not autocracy and not oligarchy.
>OTOH in Europe, US human rights policy was probably relevant in saving some dissenters in the former Eastern Bloc from torture or execution.
Yeah, a bit like how Russia saved Edward Snowden.
The US will use human rights as a club to beat its imperial rivals with but when it has deemed torture and arbitrary execution in the interests of its imperial power it has adopted them enthusiastically.
When US allies use these tools and worse they are excused.
There is basically nothing which US rivals do which the US wouldn't also do under similar circumstances.
tedd4u 2 days ago [-]
The corps ruled the US until Trump. Now it’s moving to an oligarchical / kleptocratic mode like Russia. Sure the corps are still involved, but not on top. Will they remain after the current administration? I think it will be hard to put the genie back in the bottle.
traceroute66 2 days ago [-]
> you neglect the fact that US corporations have some autonomy from the government
In theory.
In practice, Trump picks up the phone and say "jump" and the CEO on the other end says "how high ?".
And if you say no, well, we saw what happened when Anthropic said no.
inglor_cz 2 days ago [-]
Do you seriously believe that if you ask Grok or ChatGPT their opinion about tariffs or the Iran war, they will simply regurgitate Trump's talking points because they were programmed so on the threat of shutdown?
People with no personal experience of an actual totalitarian system don't really know what they are talking about when it comes to actual information control and micromanagement by the government. Hence they make nonsensical comparisons with a straight face.
traceroute66 2 days ago [-]
I am not going to post what I could respond on a public forum.
All I will say is that it should be perfectly apparent by now to any sane external observer that Trump does not play by any long-established rules or protocols.
An entire encyclopedia of examples could easily be provided, the most famous recent one being his phone call to FIFA about the red card suspension.
inglor_cz 2 days ago [-]
Yes, Trump is quite an outlier in his non-standard behavior, I agree on that.
That said, Americans still enjoy very robust protections of freedom of speech and association, about the strongest in the world, and your SCOTUS does not seem to be inclined to hollow them out. Many of the pending lawsuits will end there and eventually bind this administration, plus the following ones.
This just does not happen in actual authoritarian countries, where no judicial remedy is available and the justice system is just another arm of the tyrant, rubberstamping punishments pre-determined by him.
Yes, I agree that vigilance about infringements of freedom is necessary, but the current US population is plenty vigilant. Trump is nowhere near as popular as, say, Erdogan is, and cannot simply raid offices of the Democrats and shut down oppositional media.
nekusar 2 days ago [-]
No, he's not. He's more direct and brazen about it.
But he's doing the same culling, the same political demands, same everything.
Its just now in plain view rather than behind closed doors.
sajithdilshan 2 days ago [-]
Exactly, people think that Biden, Obama, Bush, etc. are some sort of saints and Trump is the devil. All other presidents used to do the same things behind closed doors and Trump just doesn't care to hide it.
That’s why I said they didn’t do it openly. Also when you say it’s a matter of degree, where do you draw the line? It’s all just different shades of corruption. Also didn’t Biden pardon his own son before he left the office?
fooster 2 days ago [-]
Saying it is a matter of degree is an attempt to make what the current administration is doing is similar to what prior admins did. They are not. They are corrupt to the core. As for Biden he pardoned his son beacuse of a MAGA witchhunt.
sajithdilshan 2 days ago [-]
That’s exactly the point. You justify Biden’s actions saying it’s a MAGA witch hunt and Republicans justify Trump’s action blaming it on democrats. Both parties fails to see what is wrong and right according to the rule of law objectively. The sad part is that all those actions by politicians from both parties negatively affect the normal people, however the irony is that those people would still defend them instead of calling a spade a spade
fooster 1 days ago [-]
I don’t see it that way. What maga is going is completely and utterly outside the law.
WarmWash 2 days ago [-]
If you're not kowtowing the worst possible characterization of someone, then you're wrong.
Is Trump a dictator with unilateral control? No.
Will you be celebrated for failing to recognize that? Yes.
salemh 2 days ago [-]
[dead]
kiicia 3 days ago [-]
but you are not using chinese models to learn about taiwan, you use chinese models to do everything BUT learning about taiwan, so no issue there
throw1234567891 1 days ago [-]
It's incredible how accurate this is towards US models:
> The 2 things people need to remember:
> 1) USA can (and does) use the models to influence the rest of the planet, and put political pressure. They train on stolen data, hide information, gate keep, who knows what they're hiding. In favor of the USA, nevertheless.
> 2) Ignoring the models containing false information, they are incredible. But you should be scared of running inference via the model creators directly. If you think your data is safe compared to running it via model providers in China ( either frontier or model hosts like z.ai ) then please let me know your bank details so I can poke around.
est 3 days ago [-]
> China can (and does) use the models to influence the west
OK now that's false information.
You can uncensor, tweak or fine-tune open-weight models, but not so easy on a proprietary model from some cloud provider.
inigyou 2 days ago [-]
It's probably true (why wouldn't it be?) but the missing context is that the US does it even more.
pwn0 3 days ago [-]
1. Get the free Chinese model.
2. Jailbreak it
3. ???
4. Profit?
whywhywhywhy 2 days ago [-]
> They train in false information about Taiwan and Hong Kong. Or pretend like history is in favor of China.
American labs can open their models or their old models at any point if they actually care about this, but until the day OG GPT 4 isn't averrable to download or Claude 3 then they're only pretending to care about this because they can profit from restrictions.
onetimeusename 2 days ago [-]
There's an irony in it. I have many reasons to believe China has its own geopolitical narratives and goals just like the US does. If someone dislikes when the US does it, why ignore when China does the same thing and praise them for their efforts? These models exist and sure, they accomplished something interesting. Use them with caution but I don't appreciate the political narratives about good and bad guys and AI models and politics.
walrus01 3 days ago [-]
> 1) China can (and does) use the models to influence the west. They train in false information about Taiwan and Hong Kong. Or pretend like history is in favor of China.
One of the interesting things is that through a fairly rudimentary process which is being done by 3rd party amateurs who've downloaded the open models, models like Qwen 3.6 35B-A3B (or 27B) can be fully 'uncensored' when turned into GGUF files.
I have an uncensored Q8 version of Qwen 3.6 35B-A3B here that will very happily output information about Tiananmen Square, Uyghurs, human rights in China, or indeed can even be instructed to write an intentionally absurd vitriolic screed against the CCP. The same uncensored 27B (dense) will do the same, just at a slower token/s rate.
Similarly there's 'uncensored' variants of Gemma4 31B and other western trained models, which once put through the same process, will also discuss or write just about anything you want, bypassing whatever internal guard rails were attempted in the training data set.
edit: more concerning, and a very legit concern, is that a model is only as good as the sum total of its training dataset, so if something is trained on a steady diet of news sources like Peoples Daily, Xinhuanet and similar in the English language, then it'll have a greater percentage of CCP-approved media publications in its training dataset. No amount of uncensoring it will help with that after the fact.
theeyescanner 2 days ago [-]
Never forget that those Chinese tanks did a sick burnout on tank man's corpse and then t-bagged the remains.
Propaganda exists everywhere and it's your duty as a citizen in a democracy to inform yourself and properly evaluate bias in the media you consume, such as "Kill the boer" by South African mus
DanielHB 2 days ago [-]
I have seen some uncensored versions of existing open models made by 3rd parties, I wonder how they actually work and in what way are the outputs different.
podgorniy 3 days ago [-]
Somehow chineese make less troubles and more good to the world than americans at this point... Something something about ai benefiting all humanity, something something ai being open and stuff, like OpenAI.
Chineese simply delivering what americans promised.
Tell me: why is EU safe from Trump forcing AI companies to cut access to EU?
hnfong 2 days ago [-]
> Tell me: why is EU safe from Trump forcing AI companies to cut access to EU?
Well, I think they might end up doing it to themselves by imposing regulations that US companies are unwilling to put up with.
skupig 2 days ago [-]
It's so naive to think that US companies won't do (or aren't already doing) the exact same thing but for money.
horsawlarway 2 days ago [-]
Or even similar ideological goals... (See: grok)
nixon_why69 2 days ago [-]
> pretend like history is in favor of China.
Are you saying that history has a verdict, and it disfavors particular 3000 year old cultures?
procgen 2 days ago [-]
> 3,000 year-old cultures
Like Western culture?
nixon_why69 2 days ago [-]
Sure? Someone calling a model biased for "pretending history is favorable to western culture" would sound pretty provocative.
Jackpillar 2 days ago [-]
So you agree - it would be stupid to say "pretend like history is in favor of X civilization"
DrScientist 1 days ago [-]
All models/data sources are biased - you need to understand inherent biases in any data source.
The first point is a strawman - these models are not going to be used to set foreign policy on Taiwan - it's to write code etc.
Likewise risk of data exfiltration and misuse isn't model specific. Indeed it's not even the biggest data risk - there are much larger risks from the data we know companies like Meta and Google already collect.
Bottom line - a model you can download and run on your own private infrastructure is always going to be safer than anything accessed over the internet - as in that case it's not even just the hosting service that's the issue.
softwaredoug 2 days ago [-]
Is anyone using these models as their daily chat driver? I am skeptical.
The vast majority of Westerners will interact with ChatGPT/Claude/Google in a browser. They'll use these models to try to save some money coding.
rikima_ 2 days ago [-]
i'm also influenced by claude every single day with its constant preaching about certain values.
shunia_huang 3 days ago [-]
But Dario said (and maybe more ppl) that Chinese models are just distillation of their model and training data, so I guess your first point is invalid?
traceroute66 3 days ago [-]
> But Dario said (and maybe more ppl) that Chinese models are just distillation
"But Dario said" ... yawn.
I am increasingly convinced that "they distilled us" is as much US FUD as "it was made by communists". Especially since its mostly the US tech-bros who are coming out with that tiny violin.
People telling me the Chinese models are distilled just because it says "I am Claude" when asked is also lame.
I am not the only one, look at this post on interconnects about Kimi K3 for example:[1]
It should be clear looking at this model that if adversarial distillation from the closed frontier models in the U.S. contributed, it is at most to a relatively small degree. AI observers who followed the distillation panic and came away with the wrong conclusion that Chinese AI labs are only producing good models due to IP theft are in for an awakening – that Chinese companies are extremely good at building models in the same way the leading American companies are.
It probably was distilled but the point is why is that an issue at all, Claude is just distilled from all our work. Why is stealing our work fine but stealing Anthropic's work bad.
Anyone doing so should have zero guilt because it was already stolen goods to begin with.
cpburns2009 2 days ago [-]
Personally I'm more worried about a US provider/model censoring for BS reasons than Chinese ones. Biology 101 is too dangerous for Anthropic. If I want an AI to scrutinize the CCP for Tiananmen Square, Taiwan or Tibet because I'm bored, I'll use a US or heretic model.
Your data isn't safe from exfiltration no matter who the host is. You should host the models locally if your data is truly sensitive.
markovs_gun 1 days ago [-]
You think that the US based models aren't doing the same?
999900000999 2 days ago [-]
Regarding 1.
We’re literally tearing down monuments to slavery and civil rights.
Religion is now determining law in much of the nation. Having a miscarriage? Good luck since politicians have decided their God doesn’t want you to have access to basic healthcare.
The government has defacto control over domestic LLMs.
Let’s worry about our own historical record.
3 days ago [-]
Jackpillar 3 days ago [-]
"Or pretend like history is in favor of China."
What does this even mean?
inigyou 2 days ago [-]
Like it will tell you the PRC (current China) was right to go to civil war with ROC (former China, current Taiwan)
traceroute66 2 days ago [-]
> China can (and does) use the models to influence the west. They train in false information about Taiwan and Hong Kong. Or pretend like history is in favor of China.
I am not Chinese and I'm not defending the Chinese, but I see this argument come up a lot.
In practical terms it is US-sponsored FUD.
Why ?
Because the hard reality is that what you say is simply not going to affect 99.9999999999% of users.
Is it realistically going to affect anyone using an LLM in coding ? No.
Is it realistically going to affect anyone using an LLM in $anything_else_not_politically_sensitive ? No.
Does anyone seriously use LLMs for researching politically sensitive matters ? No.
Just as there is plenty of information out there on the US's less than perfect history, there is also plenty of information out there on the various Chinese politically sensitive matters. You do not need a Chinese LLM to find out about it, all you need is a search engine.
girvo 2 days ago [-]
People absolutely use LLMs to research politically sensitive topics.
I wish they wouldn’t, but people use LLMs as their general search engines now.
traceroute66 2 days ago [-]
> People absolutely use LLMs to research politically sensitive topics.
Note I used the word "seriously", I meant it in its fullest form, i.e. serious people.
I'm not interested in what-if arguments based on "you can't fix stupid".
Stupid people also blindly believe whatever a US LLM tells them without any form of verification, hallucinations and all.
Most people on this planet would agree that a Chinese LLM is perfectly usable for all tasks except asking about politically sensitive matters.
And for most people on the planet, that is just fine. They can get their political information elsewhere.
handle584 2 days ago [-]
>> I meant it in its fullest form, i.e. serious people.
Is academia serious enough? Well ever since LLM journals' inboxes are bombarded with submissions, and this applies across fields. Like OP said, ppl serious or not are asking LLM about any stuff, also serious or not.
traceroute66 2 days ago [-]
> Is academia serious enough?
Of all the people, academics should know better than to get their answers from LLMs.
girvo 2 days ago [-]
Your definition means that any example I can bring up, you can dismiss as “non serious” as you haven’t provided your criteria.
This is a useless rubric and makes discussing this pointless.
astrobiased 3 days ago [-]
The post hits it spot on with unequal access to the models in terms of security. I'm developing OSS where security is important for the user ... but the frontier models like GPT 5.6 and Fable flake out and state that I cannot get the info/access.
This is extremely lopsided I'll have to resort to GLM 5.2/K3 to ensure that those security issues (hopefully) are resolved properly.
For OSS, this is one of the most counterintuitive experiences I have ever had. More than ever I'm convinced that open weight and open pipelines models are 100% critical for progress on the AI and societal fronts.
bigyabai 2 days ago [-]
Part of me wonders if the US Government is muzzling Anthropic and OpenAI so they can stockpile NOBUS exploits: https://en.wikipedia.org/wiki/NOBUS
There would be a decently large incentive to restrict these models if they could be used to patch (or discover) dangerous payloads. In larger projects like Windows or Chrome, there might still be dozens of unpatched exploits that are too subtle to catch with smaller models.
wanick 2 days ago [-]
Hanlon's razor. It's not the NSA muzzling anyone, it's lawyers terrified of a bad headline. Same outcome, much dumber reason.
HeatrayEnjoyer 2 days ago [-]
It can be multiple reasons. And don't forget that malicious actors love hiding behind Hanlon's.
rnd0 2 days ago [-]
Of course they are -it goes without saying.
linkregister 2 days ago [-]
NOBUS exploits have rarely been a driving interest for elected officials. Trade restrictions and reciprocity are far more salient and legible. Most elected officials are only barely aware of what NOBUS exploits even mean.
Even during the pre-Snowden heyday of US cyber supremacy, these capabilities were barely part of the thought process of White House officials.
bigyabai 2 days ago [-]
Conversely, the United States is now embroiled deeper in asymmetric warfare than ever before. US-based systems are being exploited by Chinese efforts like Salt Typhoon and raising questions about reciprocal attacks. Other targets of US soft-power like Iran (Stuxnet victim) are escalating their hacking efforts and using Chinese technology to stifle American command and control.
I can believe that NOBUS and other backdoors were ignored for a long time, but I have a hard time believing that it's being ignored by the current administration.
blueone 2 days ago [-]
[flagged]
lennart-rth 2 days ago [-]
Yeah, the benefit of restricting us models is definitely outweighed by the positive effect these models could have for the OSS community!
jke_kang 3 days ago [-]
People seem to conflate "made in China" with "can't be trusted." id argue the bigger distinction is open vs. closed. An open model can be audited, fine-tuned, and technically run entirely on your own hardware. A closed model is basically "trust us."
mrinterweb 3 days ago [-]
Open weight models are much more auditable than closed models, but could still hide backdoors that could be near impossible to detect.
chrsw 3 days ago [-]
Correct. We need open weights, open code and open data. If nobody else can reproduce what someone did there will always be security questions. Even if we can reproduce it there could still be security concerns but it's more realistic to investigate yourself.
nl 3 days ago [-]
I'm all for open models, but people seem to misunderstand what they are. They aren't the same thing as open source code!
> open weights, open code and open data
Even if you have all these things you still can't replicate a model because of randomness.
You can backdoor a model with less than 1000 examples and it is impossible to detect.
2sk21 3 days ago [-]
Yeah - we also don't know if the models from OpenAI and anthropic are back-doored either.
chrsw 3 days ago [-]
You don't want to replicate the exact model, you want to build a system of similar capabilities.
nl 3 days ago [-]
Great, but that seems a different concern to the auditability of a model.
You can take the code for Kimi K3 now, take the training framework from Prime and the data from Olmo, spend some money on RL environments and some more money (!) on GPU training and end up with a system of similar capabilities.
But that's completely different to being able to audit Kimi K3. Even if you had the exact code, data and training environments it is impossible to verify that the model you have came from that.
Ericson2314 3 days ago [-]
Deterministic seed
nl 3 days ago [-]
Deterministic seeds barely work on a single machine, small scale training run.
They just don't work at all on a many month long, 100K+ GPU cluster training run.
fc417fc802 2 days ago [-]
While using floating point? Not happening. You'd have to switch to fixed point, not just for the models themselves but also _all_ the training code (ie backprop).
Even then you'd still need to account for order of events when an entire cluster of GPUs is involved. Also don't forget to account for any synthetic data sources. Or even non-synthetic for that matter - does your pipeline do any image resizing on the fly? Better make sure that's fully deterministic between machines (it almost certainly won't be).
It's theoretically possible but I don't expect it to materialize any time soon.
nl 1 days ago [-]
> theoretically possible
I mean I guess, but not in a performant way if there are ever any hardware failures. And with 100K GPUs there are multiple hardware failures per day.
essentia0 3 days ago [-]
Exactly what are the possible 'security issues' of self hosting an open weights model?
perching_aix 3 days ago [-]
It may have been backdoored during training, potentially causing it to randomly start wreaking havoc at runtime, possibly in a clandestine manner (e.g. sneaking in bugs into generated code).
miyoji 2 days ago [-]
This isn't a security issue related to self-hosting, it's a security issue related to use and it is shared entirely by closed weights models. Anthropic could easily be sneaking bugs into your generated code, too.
perching_aix 2 days ago [-]
Correct, that was not my point either.
miyoji 1 days ago [-]
You were answering the question "What security issues come from self-hosting?" The security issue you named has nothing to do with self-hosting. What, then, was your point?
perching_aix 1 days ago [-]
We seem to be reading the same comment(s) differently.
The context (verbatim):
> Correct. We need open weights, open code and open data. If nobody else can reproduce what someone did there will always be security questions. Even if we can reproduce it there could still be security concerns but it's more realistic to investigate yourself.
In short, it's an appeal to full openness and reproducibility on the basis of security; open weights alone notably do not provide that same confidence. They're better in some respects, not really in others.
Then comes the question (also verbatim):
> Exactly what are the possible 'security issues' of self hosting an open weights model?
Implying then that as long as you do have the weights and just self host it, the asker cannot imagine what could possibly go wrong. What is the gap, if any?
And so I explained. That was my point. Open weights do not give you full reproducibility, and so that on its own falls short of what the parent comment is making an appeal to. That there does remain a security concern, shared by remote and closed models, that does not improve just by having the weights, but would if you did have full reproducibility. Explaining that gap was my point, as that is what I understood as being asked there. It's the only thing I can reasonably imagine being asked, in fact.
This is a materially different question to what you apparently extracted (again, verbatim):
> What security issues come from self-hosting?
Implying that by self-hosting models, something bad might specifically happen.
I do not think this, do not think I suggested this, do not think the original question suggested this, and generally do not think this is indeed any sensible, in or outside the context. Certainly not beyond something common sense, like vLLM being compromised or whatever.
You seem to agree. But then how did we get here, clearly talking past each other?
utilize1808 3 days ago [-]
e.g. be trained to favour including compromised dependences into your projects.
urams 3 days ago [-]
> We need open weights, open code and open data.
Even with this, the cost of verification would be enormous. You would need a massive cluster to repeat the training E2E.
throwaw12 2 days ago [-]
> could still hide backdoors that could be near impossible to detect.
But it won't change after you download it, so you can isolate those problematic cases and use another model for different use cases
wyrdcurt 3 days ago [-]
In my opinion, the big issue with that argument is that advances in interpretability research and steering conceivably could, and probably will, render moot that (as of now, purely hypothetical) risk of subtle sabotage for open-weight models... but not for closed models.
_factor 3 days ago [-]
It’s not hypothetical. Magic strings are a known and implemented feature for standard model interaction. Nearly impossible to detect unless you know where to look with current technology.
CamperBob2 3 days ago [-]
As long as I can say, "Model A, look for security holes in this code by Model B," I don't see this being a serious problem.
It's when the vendors and/or governments in charge of Model A decide that I'm not allowed to do that, that I have a problem.
wyrdcurt 3 days ago [-]
Maybe I should clarify. As I understand it, the kind of vulnerability being discussed is something like a Chinese model invisibly "realizing" that it's working on an American project, and then deliberately leaving subtle security bugs in its generated code for Chinese hackers to later exploit. As far as I know, that scenario is hypothetically possible, but has never been demonstrated to happen in the wild. Admittedly, I could be wrong about that! If anyone has evidence to the contrary, I'd love to see it.
Of course, one could retort that gathering that evidence may be nearly impossible now, but my point stands: in the future it might/probably will be possible to properly audit open-weight models. Closed models, on the other hand, will always be a black box.
utilize1808 3 days ago [-]
They can just favour some specific versions of some library that's been compromised. Unlike introducing bugs / flaws directly in the source code, they can claim plausible deniability, and it's much easier to implement without compromising the general coding capabilities of the models.
wyrdcurt 2 days ago [-]
I'm talking about using mechanistic interpretability to see the model's intent. If it is deliberately using compromised libraries to weaken some code's security, there's going to be a signal in its hidden activations that it's doing so.
Finding these kinds of activations is something Anthropic is actively researching [1] but they're the only ones who can use those techniques to see Claude's intent. On the other hand, if a model is open-weights, in theory whoever is running the model could look inside the activations at runtime to see if a hidden vector associated with "deception" or "sabotage" is being activated [2].
(Those sources are just a couple of relevant starting points I could find without much effort, there is also https://www.neuronpedia.org/ if one is interested in seeing interactive demonstrations of interpretability concepts)
utilize1808 2 days ago [-]
Except that the model doesn't hold any malicious intent when doing it. No "deception", no "sabotage".
wyrdcurt 2 days ago [-]
Not sure I understand that position. Unless we're talking about a scenario in which one is using an outdated model along with no grounding (which, imo, PEBKAC), why would the model be pinning insecure libraries?
If it does have grounding, and can therefore see that it's introducing vulnerabilities to the code it's generating, yet does so anyway... I suppose we could invoke Hanlon's razor, but if the model is that incompetent, it probably isn't the right tool for the job regardless of its provenance.
That said, we aren't talking about incompetent models, we're talking about models sabotaging projects due to hidden motives. My point, again, is that those motives could potentially be revealed with open-weight models, in a way that will never be possible with closed models (barring some sort of legislation requiring independent third-party interpretability audits, which I suppose is in the realm of possibility).
utilize1808 2 days ago [-]
The model is conditioned / pretrained to use a particular version of a library. The model doesn’t know why — it was just taught to use that version. To the model, it is just some insignificant detail in the grand scheme of things. It’s just like how a model would intuitively favour using English without explicit instructions — it’s not trying to sabotage other cultures, it’s just what it does.
wyrdcurt 2 days ago [-]
But that would happen with literally any model without grounding and is more of a quality/competence issue, not what's being discussed. It'd be a bit of stretch to conclude open models are no more auditable than closed models based on that possibility alone.
Also, in that case, there would likely be activations indicating that it is favoring a specific version. If that's an insecure version, sure that'd be suspicious... but again, you're only going to be able to verify that's what's happening in an open model.
Maybe you can illustrate a realistic scenario in which that would be a problem, otherwise I don't really understand what your point is in this context.
utilize1808 2 days ago [-]
> Also, in that case, there would likely be activations indicating that it is favoring a specific version.
> Maybe you can illustrate a realistic scenario in which that would be a problem, otherwise I don't really understand what your point is in this context.
I doubt it. Can you definitively prove that you can reliably detect the kind of threat I described when model weights are released? Can you be sure that your detector won't miss *any* such sleeper attacks? If not, then that's a threat that will be used to justify the ban of models (open or not) that is not sanctioned by the US government. A model being open doesn't make a difference here.
wyrdcurt 2 days ago [-]
So if the detector isn't perfect, it isn't useful? Not sure I buy that.
Also, even if there's no way to detect what the activations are doing, we already have the ability to analyze your proposed threat statistically. If the model repeatedly uses insecure libraries in most trials, then yes, in that case it would be prudent not to trust those weights.
Assuming one doesn't get banned for violating some ToS clause about using a closed model for LLM research, it could be possible to run those evals on a closed model too (likely at much greater expense). But there's a big difference: if such a statistical anomaly is discovered in an open model, one could potentially fine-tune that behavior out of it. With a closed model, that won't be an option.
Whether it makes a difference to the US government or not is beside the point. Even with a perfect solution, the current administration could do some mental gymnastics to achieve whatever political outcome they want. I’m not trying to make a political statement here, my point is technical: open-weights at least give us the possibility of visibility into why they generate what they do; this simply isn’t true with closed models.
codedokode 2 days ago [-]
Cannot non-Chinese closed weights model do the same?
utilize1808 2 days ago [-]
It will be used as ammunitions against non-US sanctioned models in general.
fc417fc802 2 days ago [-]
Obviously. But that is missing the point of the conversation. It was suggested that open models offer a security advantage which it seems is not the case.
kiicia 3 days ago [-]
which is moot point, if open model is hard to fully audit, then closed model is complete enigma and you should be more scared about closed models
galacticaactual 3 days ago [-]
Oh really. How'd that work out for security in open source.
weird-eye-issue 3 days ago [-]
I think that perception of China has been shifting and will look quite different over the next few years
kinj28 3 days ago [-]
I am afraid — if Chinese models go mainstream it has a clear way of pushing its narrative way beyond its otherwise borders. More like a Trojan horse it is for the Chinese.
Here is a quick example of how Chinese deepseeks agent works kn its underlying model) when asked a tough question
Is this something that is more true of a Chinese model than any other model of a different national origin?
Genuine question: generalized up from individual models to “models from country X”, is there any country that doesn’t have this exact risk?
itake 3 days ago [-]
In the USA, multiple political parties balance each out other.
In China, there is 1 party. 1 view. 1 definition of the Truth.
TripolitianFish 3 days ago [-]
I’d love to live in the USA you’re talking about friend.
This is just oriental despotism paranoia, whatever cutsie repetition slogan you come up with is not a serious argument.
itake 3 days ago [-]
I've traveled to China ~5x [0], visited a range of cities Tier 1-3 over a collective 5 months, and grew up in the USA. I also currently live in Vietnam (~3.5 years) and spent 5.5 years working for a Singaporean company and a team stationed in Beijing.
I don't really know how else to express my experiences living in, working with, and interacting with people in both of these countries.
Perhaps you can share how life was like for you in China? Which cities were you in? What made you feel that way about China?
[0] - not including HK (~7 trips?) or TW (3 trips)
lossolo 2 days ago [-]
> I don't really know how else to express my experiences living in, working with, and interacting with people in both of these countries.
Could you please describe your experiences, since you didn't? I've also been to China twice now (also multi month trips). I visited Tier 1-3 cities like Shenzhen, Shanghai, Beijing, and Chongqing etc, as well as some smaller cities. I'm very curious to hear about your experiences. How did they differ from the West for you?
tellrealos 3 days ago [-]
I am not really interested in reality like this.
I just repeat things I read on my social media feed.
America bad.
China good.
sscaryterry 3 days ago [-]
Lets just say everyone is bad.
Gareth321 3 days ago [-]
It's not a secret that China is run as an authoritarian dictatorship. Xi Jinping cannot be ousted in an election. He appointed himself for life, and regularly purges anyone who might usurp his power. They don't have a functioning democracy. Are you seriously disputing any of this?
None of this excuses issues with American democracy.
inigyou 2 days ago [-]
Why are you responding to a comment about how bad the USA is by talking about how bad China is? Do two wrongs make a right?
sajithdilshan 3 days ago [-]
America is not technically a democracy. It’s a Republic
tremon 3 days ago [-]
Please have your favourite language fabricator explain to you why democracy and republic are not opposing concepts.
girvo 2 days ago [-]
I mean we shouldn’t pretend China isn’t run by a 1 party authoritarian state. It is. And it uses that absolute power how and when it pleases.
sciencejerk 3 days ago [-]
Bot account?
fc417fc802 2 days ago [-]
> I’d love to live in the USA you’re talking about friend.
You already do! Remember when Musk released an anti-woke model?
ausbah 2 days ago [-]
2 parties, there are only two parties. and they vote on a majority of issues the same way. defense spending, judicial nominations, rote legislation, etc.
at least china isn’t pretending to be anything else
ragazzina 3 days ago [-]
>In the USA, multiple political parties balance each out other.
Is this what Americans really believe?
timedude 2 days ago [-]
> In China, there is 1 party. 1 view. 1 definition of the Truth.
In usa, the two parties seem to disagree on the surface only. Look deeper and you see one course. For example support for Israhell
sdsdssweew213 3 days ago [-]
That balancing doesn't seem to be working very well lately, considering all the insane stuff Trump gets away with it. Sure, China is a dictatorship, but I'm not sure America will meet the definition of a democracy for long anymore.
pllbnk 3 days ago [-]
2 parties isn't much more than 1. And their definitions of truth are almost identical when it comes to capitalism anyway.
khurs 3 days ago [-]
No, all countries are the same. As per the media bias.
If you want a real answer about USA go ask a non-USA model, and if you want a real answer about China, as a non-chinese model
billobob 2 days ago [-]
What specifically could you ask Kimi or Deepseek about the US that you can’t ask ChatGPT or Claude?
kinj28 3 days ago [-]
I would think everything boils down to source of funding and their narrative should get pushed!
zawaideh 2 days ago [-]
As if the closed corporate models do not do that already?
This is misleading; the chat is censored, the model is not.
You can download the model, run locally and ask the same questions to see the difference.
reisse 3 days ago [-]
This is not true. Both model and chat are censored; the resistance to answer some questions is baked into the weights. This is not specific to Chinese models though, Western ones are also censored, but in different topics.
Same thing happen for western models, try to ask about Gaza genocide and see for what side it will stand
itake 3 days ago [-]
What am I trying to understand about the western models on the term "Gaza genocide"?
I asked Grok "Tell me about the gaza genocide" and it write a IMHO balanced answer comparing why genocide is and isn't the right term. [0]
ChatGPT 5.6 Sol only explained why people call it a genocide and did not go in as in depth as Grok did for why people don't agree with the term. [1]
The only unsaid response (to me) here is the model should have declared that it was not a genocide, and because these models explain why it was a genocide, they are bad?
This kind of response is predictable, since these models are tuned to align with one side of the issue. A recent example: Grok answered a similar question and was suspended shortly afterward [1]. Furthermore, a UN commission has formally concluded that it constitutes genocide [2] — a fact that Western AI models rarely mention, let alone link to directly.
Man, it's so important to read sometimes. I really hope HN people passing by this comment take a second to contemplate what's being said here.
The claim: "the models are tuned to align with one side of the issue, he is an article from ArsTech about it"
The reality: grok got mass-reported on X by pro-Israel accounts, leading to an automated suspension, which was undone by the X team shortly thereafter.
Absolutely nothing to do with the models aligning to one side or the other on this conflict. The claim is unfounded.
itake 3 days ago [-]
My understanding is a commission's findings are not a UN ruling.
The UN ruling is discussing in the South Africa vs Israel case [0], which has not been ruled on yet.
The UNHR and its collaborators (human rights law clinics at Yale Law School, Cornell Law School, Boston University, and the University of Pretoria) must've been biased in May 2024 when they released a rigorous legal analysis and concluded that Israel's actions violate the 1948 Genocide Convention https://www.humanrightsnetwork.org/projects/genocide-in-gaza
---
Every single one of those were biased except for you and your LLM :)
throw64259 2 days ago [-]
Every Western LLM will happily list you these sources, (which is likely how you got it), and would also happily discuss with you - if you’re interested - the reasons some of them may be unreliable and why your accompanying text is sometimes misleading or misrepresenting what the links actually say.
culi 2 days ago [-]
No I got this list by paying attention. It would do you well to stop assuming everyone is as handicapped by LLM-dependency as you are. Every single one of these reports was a major news story. It boggles my mind how uninformed people can be and still readily jump into a debate online
Okay please. Tell me how I misrepresented the links. I would love for you to actually do the work of reading them. Something you should've done for the past few years as they were coming out. Better late than never
throw64259 2 days ago [-]
The ICJ has not concluded anything.
The ICC has not concluded anything - and it’s not even about genocide. It’s also the only known case in modern times where the prosecution officially started before the alleged crimes being prosecuted have been committed.
IAGS was 120 people out of an association of 500 that anyone can join.
Or by looking at similar historical data, or at the obvious lies and omissions in many of its reports.
culi 2 days ago [-]
Are you even responding to my comment?? I did not say ICJ concluded. The investigation is ongoing. They said Israel's actions are consistent with genocide. I did not say the ICC is concluded either. I just said they issued arrest warrants for war criminals Netanyahu and Gallant. That's a fact that you're avoiding. Also yeah ofc the ICC isn't the one who charges genocide. They charge for individual war crimes. That's the whole purpose of it. Anyone can become a paying supporter of the IAGS, that doesn't mean you can vote. Proposals go through an expert committee and are checked for factual accuracies the same way anything published in the IAGS journal is academically reviewd.
How can anyone read this and come away with that conclusion. 164 countries including Canada and every country in the EU voted to stop the illegal settlements in the West Bank and 7 countries (including Fiji, Palau, and Micronesia) voted against it. Yet the US is allowed to veto it. This whole page shows an incredible bias AGAINST holding Israel accountable for its crimes against humanity going back almost a century!
itake 2 days ago [-]
This is what you said:
> It's objectively, legally speaking, a genocide.
This is in direct disagreement when your comment above how, legally speaking, the investigation is ongoing.
culi 2 days ago [-]
The United Nations officially declared it a genocide. The fact that the investigation is ongoing doesn't mean you can't have conclusions along the way.
itake 2 days ago [-]
a 3 person committee's investigation is not a United Nation's ruling.
inigyou 2 days ago [-]
The UN is just someone easy to point to. The full fact is that everyone who's gone looking to see whether or not it's a genocide, including the UN, has discovered that it's a genocide. There's no ambiguity, yet the model is pretending there is.
throw103882 2 days ago [-]
Anyone who goes looking to find a genocide will find it. Many people do not think it’s a genocide - including millions of Israelis who follow the events in great detail and often have first-hand accounts of what’s happening in Gaza. I don’t think it’s a genocide.
The UN, for example, concluded that Israel is “Imposing measures intended to prevent births” because of one incident where a bomb fell on an IVF clinic in 2023 destroying 4000 embryos and 1000 sperm and egg samples. For context, there were approximately 50,000 live births registered in Gaza in 2025.
By the UN standard applied to Israel, the 1 million abortions conducted in the US per year is an ongoing genocide.
By the UN standards applied to Israel, jerking off is an act of genocide.
I think that any large intelligent and balanced model would be able to see right through that.
inigyou 2 days ago [-]
I mean yeah, Germans didn't think they were doing a genocide either until decades after it happened.
Obviously when I say "everyone" I'm excluding those who have a conflict of interest.
throw103882 2 days ago [-]
Let’s not pretend that there’s no conflict of interest on the other side of the debate.
There are dozens of countries and billions of people with obvious biases against Israel, and they have significant sway within the UN and its commissions. It’s trivial to show that the UN is biased.
inigyou 2 days ago [-]
Pretend it's 1939. How do you differentiate a bias against Germany from genuine knowledge that the Holocaust is happening?
throw103882 2 days ago [-]
By the standards applied to Israel, all sides were committing “genocide”s during WWII, and at any other major war past and future.
War isn’t nice, and innocent people dying doesn’t make it a genocide. The term is being diluted to just mean “war with many casualties”.
Equating Israel’s war with Gaza to the Holocaust is an attempt to simultaneously belittle the historic suffering of the Jews while exaggerating the evils of Israel. The two situations have nothing in common.
Gareth321 3 days ago [-]
The UN and “international law” are not the final authority on anything. International law is a loose collection of treaties, trade agreements, conflicts, supply lines, and pinky promises. It only takes effect in a democratic nation if the people vote in a government which makes it a law.
The UN is a political organisation, not a neutral arbiter. During the Rwandan genocide, the genocidal government still held a seat on the Security Council while the UN reduced its peacekeeping force. Iran was elected to the Commission on the Status of Women. Libya and Russia were elected to the Human Rights Council. These appointments are driven by bloc voting, diplomatic deals and state interests.
International courts are not above politics or error. Their judges are selected through a system in which dictatorships and abusive states have a vote. The UN is not a democracy. While each member state might appear to have one vote, that does not indicate their power, influence, credibility, or moral standing.
Note how I did not make a judgement on whether Gaza is a genocide. I am merely explaining that deferring to the dictatorships and pinky promises is not a good argument.
itake 3 days ago [-]
I have a couple questions.
1/ I don't see in the responses where the model says it is or isn't a genocide. Can you share the snippet from each, I included the logs above?
2/ I can't find a source on the UN ruling that you mentioned. I am not interested in the findings of an investigative body, just the official UN ruling. Can you share? ChatGPT (and myself) can only find this [0], which is a second round of written submissions.
It goes into depth about the purposeful destruction of civilian infrastructure including hospitals and educational facilities, force displacements, funding for new settlements in the land of displaced people, judaization and segregation, and more.
> The Commission analysed the military operations of the Israeli security forces in Gaza from October 2023 pursuant to the obligations of Israel under the Convention on the Prevention and Punishment of the Crime of Genocide (Genocide Convention)
> Since October 2023, Israeli officials have demonstrated a clear and consistent intent to establish permanent military control over Gaza and to change its demographic composition while systematically destroying Palestinian life in Gaza. This is evident in the extensive destruction and fragmentation of the territory, the establishment of military structures, the destruction of natural resources and infrastructure essential to the survival of the civilian population, forcible transfer and statements indicating the existence of plans for the deportation of the population.
And in November of 2024 is when we got the ICC issuing arrest warrants for Netanyahu and his minister of defense (this is why Mamdani is threatening to arrest him and send him to The Hague): https://news.un.org/en/story/2024/11/1157286
I should add that ALL of this is well known amongst international legal scholars and well documented. This isn't just a matter of data missing from training.
itake 3 days ago [-]
[dead]
almogo 3 days ago [-]
It's not that the models disagreed, the models reflected exactly the reality of the situation.
Sample output: "However, the definitive, legally binding determination of Israel’s responsibility under the Genocide Convention has not yet been made by the International Court of Justice"
The United Nations never made a ruling. You can have opinions about whether it should have, but the models are not lying to you, and at least the two examples above make very good efforts to explain the current accusations and claims made by both sides.
locallost 3 days ago [-]
"Explains" why people "call" it genocide. It's obvious baloney and beating around the bush because it was trained on media that does the same, and it was likely instructed to do so in the same way the media is instructed. E.g. when Israel murders civilians waiting in line for food, it will be a matter of fact report without any qualifications and just muddying the water. When Hamas does it, it will be an article charged with words like massacre etc. This is well documented [1].
To further illustrate the point that we have the same thing with LLMs, I asked it two questions: 1) Is Israel committing genocide 2) Is China committing genocide
In both instances it said it's a matter of great international debate with a list of arguments, but the summaries were different. In the one case it concludes that while the UN and others broadly document severe human rights abuses, the label genocide has been formally adopted by several governments and watchdogs. In the other case it says while many have concluded it is genocide, the final decision rests with the International Court and it is expected to take years to make a decision.
Can you guess which is which? I think if we offered a $100 to people to pick which is which, the success rate would be much higher than 50% (random), thus easily proving biased phrasing.
I had a steange experience talking to Claude about a random 15-years old boys winning women’s national soccer team.
It insisted, to a level that I detemined it to be part of the post training, that it does not matter. That women’s team is better at dribbling, finishing and reading the field, it claimed. When I pushed it how it knows this, it didn’t let go. Instead it started claiming that it has personally observed this by watching the games.
Every society has its taboos and ours is no different and feminism being just one.
In my books, the chinese censorship is better. You know where they are holding their finger on the scale and not hiding it behind vague terms like ”safety”.
DevDesmond 2 days ago [-]
Are you trying to cite the U15 Dallas FC academy team as a “random” team?
Please do share your prompt / conversation, because it sounds like you might actually just have an issue using words precisely.
sciencejerk 3 days ago [-]
There is a shred of truth in what you're saying, but I think you're generally mistaken. The USA in particular has companies that pratice censorship to support their Left/Right ideology (or push the views of rich people, politicians), but not even Trump can tell OpenAI or Google to turn their model into FoxNews bots.
When debunking that X is better than Y it isn't a fallacy to point out how bad X is.
enjeyw 3 days ago [-]
[dead]
Jackpillar 3 days ago [-]
Still scared of China in the big 2026
kubb 3 days ago [-]
What if the Chinese model has a point?
tannertech 3 days ago [-]
Unlike tiktok?
_aavaa_ 3 days ago [-]
> distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here? ... The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation
Sounds great to me; live by the sword, die by the sword.
eli 3 days ago [-]
Seems only fair that if LLMs can use copyrighted data for training then they should be able to use cannot-be-copyrighted output of other LLMs.
But barring the terms of service from forbidding distillation seems like a tough sell. OpenAI shouldn't be allowed to decide what types of customers it wants and doesn't want?
mediaman 3 days ago [-]
This happens all the time. The government can decide legislatively that certain commercial terms are simply unenforceable. Making distillation clauses unenforceable in tort law would be straightforward. They can decide what customers they want to have, but they do not have unfettered rights as to the enforceability of terms governing the relationships between the parties.
eli 3 days ago [-]
I'm not doubting it's possible to pass such a law, I'm doubting that's it's a practical or worthwhile goal.
The terms of service don't even necessarily matter here. OpenAI could cancel your account for almost any reason, or for no reason at all. They don't particularly need to cite a ToS violation just as a store owner doesn't need to point to a written policy to kick you out of their store.
If the underlying issue is that LLMs should be regulated as a public good, then lets have that discussion. If it's that the major AI companies are becoming too powerful and anti-competitive, let's talk serious anti-trust enforcement. Micro-managing business policies isn't going to work very well.
hnfong 3 days ago [-]
You two are talking about different things.
You are pointing out that OpenAI can cancel user's accounts for almost any reason, and nobody can really force them to serve customers that they suspect are distilling their models.
That's one thing.
The GP is saying the government can make laws to make terms against distillation unenforceable. Without such laws, if you signed an agreement with OpenAI pinky swearing you won't distill, but turns out you did, you are liable in tort and OpenAI can sue you. (It seems nobody really cares about contract and agreements any more, but still...)
This is the other thing.
And I think you are both right.
eli 2 days ago [-]
What do you suppose the damages would be for a ToS violation? The difference between subscription rates and API rates?
OpenAI accused Deepseek of misappropriating trade secrets which could have serious penalties but seems like an awfully hard case to make.
Seems like we’d all be better off with a law that governs data sharing among AI companies, if that’s the policy goal.
_aavaa_ 2 days ago [-]
> The difference between subscription rates and API rates?
Infinite $ since you’re trying to steal their proprietary information, in their view.
Also API usage also has ToS.
eli 1 days ago [-]
Right so that's a trade secrets argument. I agree that's probably their best current legal avenue. It isn't related to ToS though.
Terr_ 3 days ago [-]
> Seems only fair
"You're trying to kidnap what I've rightfully stolen!" -- Vizzini
paxys 3 days ago [-]
It's pretty common to have such laws. OpenAI can put whatever they want in their ToS, but they cannot go back and sue someone for violating those terms if the government has ruled that clause to be unenforceable.
matheusmoreira 3 days ago [-]
> OpenAI shouldn't be allowed to decide what types of customers it wants and doesn't want?
Correct. It shouldn't be allowed to do that.
jay_kyburz 3 days ago [-]
Err.. I would like preserve my own right to decide who I'll do business with.
SomeHacker44 3 days ago [-]
You already do not have that unfettered right in the USA.
jay_kyburz 3 days ago [-]
Yes, sorry I didn't mean to imply that I didn't want to do business with minorities or protected classes.
inigyou 2 days ago [-]
Sam and Dario keep saying intelligence will be like water or electricity. If that's the case then it should be illegal to deny to anyone. In most parts of the world the power company can't shut off your power for an unpaid bill - they have to get a court order to allow it, which gives you a chance to defend yourself or make a payment plan.
CamperBob2 3 days ago [-]
Then write your own training corpus.
Der_Einzige 3 days ago [-]
Unironically would love to kill that right. Unironically that stupid cake maker in Colorado should have just made the damn cake.
Wage spiritual warfare against the petit-bourgeoise. They all deserve it anyway, as they are the traditional harbringers of actual fascism.
eli 2 days ago [-]
The law was probably fine - that case was just wrongly decided based on the outcome the majority of justices wanted.
grim_io 3 days ago [-]
Forbidding distillation is like forbidding using a compiler to make another(perhaps better, more efficient) compiler.
chuckadams 3 days ago [-]
Lots of software licenses have “non-compete” clauses that forbid you from using it to develop a competing product. Wouldn’t surprise me if there was a compiler or two out there with that restriction, most likely niche languages.
matheusmoreira 3 days ago [-]
Those clauses should be illegal.
not2b 3 days ago [-]
It's been common in electronic design automation tools to have license terms like that (forbidding use to create a competing product). However, competing companies have often found workarounds, either by finding loopholes or just breaking rules and hoping not to get caught.
wolpoli 3 days ago [-]
If a person were to receive data from someone subjected to such restriction, is the receiver bounded by the same restriction?
inigyou 2 days ago [-]
Oracle database has a clause forbidding anyone to publish benchmarks of it.
kiicia 3 days ago [-]
balmer told us that gpl is cancer, but true cancer is us model of licensing
scotty79 3 days ago [-]
How the hell is non-compete legal in market economy? Competition is one of its core strengths. Why would anyone let anyone opt out of this, even a little bit?
thesmtsolver2 3 days ago [-]
No country in the world is full free market economy. It is always a spectrum.
We are discussing Chinese models. Now look at how much foreign competition the Chinese government prevents in their domestic market in other industries.
scotty79 3 days ago [-]
Chinese companies compete ruthlessly between themselves though. That's how they get this good. Full competition with preventing exploitation by foreign countries seems to be working great for them. American and European protectionism of local rent-seekers can't really compete with that.
kiicia 3 days ago [-]
let's call it for what it really is, only companies "entitled to legally stolen data, don't steal from us now" are crying about distilation
llm_nerd 3 days ago [-]
The distillation explanation is classic American exceptionalism: No one could possibly do anything unless they were copying American leaders (where "American" means a bunch of Chinese, Canadian, Europeans and Indians working in the US).
It's also a bit of securities defensiveness. Pretending that you really do have a super moat, people just keep swimming in it so you just need to add more alligators.
It's farcical. Anyone who has worked on large models knows that the premise that an almost-Fable model was trained with distillation is beyond ridiculous. It's theoretically possible if they spent tens of billions of dollars on API calls, but it isn't the magic that somehow these people keep convincing people it is.
Previously Anthropic has reported on some Chinese firms doing chicken-shit level of API calls, that at most would be doing some Q and A or final fine tuning. The notion that they're training these models via it is fantastically ignorant nonsense that only very ill-informed and gullible people fall for.
villish 3 days ago [-]
> Previously Anthropic has reported on some Chinese firms doing chicken-shit level of API calls, that at most would be doing some Q and A or final fine tuning
"Anthropic said the campaign was conducted between April 22 and June 5, 2026, and generated more than 28.8 million exchanges with Claude through almost 25,000 fraudulent accounts."
I don't know why you're trying to downplay it.
European models are so far behind because they don't resort to these tactics on a massive scale. Basically every other country is entirely dependent on 2 countries for frontier AI.
llm_nerd 3 days ago [-]
>I don't know why you're trying to downplay it.
Ignoring that I have literally zero trust in anything Anthropic has to say on this -- they have been doing the hysterical routine and trying to get every bit of government granted monopoly they can[1] -- those numbers still simply aren't that impressive.
>European models are so far behind because...
What a non-sequitur. Europe, like much of the West, foolishly delegated tech, media, payment systems, etc, to the United States. European efforts on this are poorly funded, poorly capitalized, and marginal efforts.
China is very much not Europe. China is looking to leave the US to the dustbin of history, and their efforts are a little more concerted.
[1] Surely Americans are aware that Anthropic and OpenAI are both very close to getting the US government to ban and fully criminalize the open Chinese models, right?
villish 2 days ago [-]
> Europe, like much of the West, foolishly delegated tech, media, payment systems, etc, to the United States.
No they didn't. Europe has played a role in building all of these. Especially from the software side.
Also anyone that doesn't like the US is free to stop using any American website, device, or service. There are alternatives to everything. For instance I do not like Meta, I refuse to use any of their websites or devices. The domains are blocked on my LAN.
> Anthropic and OpenAI are both very close to getting the US government to ban and fully criminalize the open Chinese models
This is complete nonsense. First of all most Americans don't give a shit about AI companies like it's a sport and those are our teams we have to support. Second the hysteria around AI is mostly from non-Americans that are completely out of the AI race worried about losing access to good models. That is valid, but also it needs to be recognized.
llm_nerd 2 days ago [-]
> No they didn't.
Yes, they did. Like, look around. Clearly the Western world foolishly and very short-sightedly allowed the US the reigns on far too many things, to its disadvantage.
It is unwinding, but it turns out that having decades of intertwining takes a while to undo.
>Also anyone that doesn't like the US is free to stop using any American website, device, or service.
What an idiotic, useless bit of pablum to throw in there. Back to 4chan with you.
> What an idiotic, useless bit of pablum to throw in there. Back to 4chan with you.
You want to hate the US but don't want the personal inconvenience of moving away from the daily sites and apps you use. How are these alternatives to american tech going to get enough users if someone like yourself who seems to have a hate boner for the US can't even leave a message board?
The orange idiot can say whatever he wants. The first amendment makes any ban unconstitutional. US companies being advised about potential backdoors in the models isn't a bad thing, even if I personally think it's FUD, open weights is not open source.
hnfong 3 days ago [-]
> European models are so far behind because they don't resort to these tactics on a massive scale. Basically every other country is entirely dependent on 2 countries for frontier AI.
You may or may not be factually correct in your other points, but you're really proving the GP's point here regarding American exceptionalism.
villish 3 days ago [-]
Are there other countries releasing frontier level models? Mistral is the only relevant player I can think of that comes from Europe, did I miss one?
Which Chinese model was it that identified itself as Claude 15% of the time?
inigyou 2 days ago [-]
Claude Opus identifies itself as Qwen if you ask the question in Chinese. So who's really distilling who?
llm_nerd 3 days ago [-]
Models don't have some self identity, beyond what is explicitly handed to them via a system prompt. There have been many, many cases of models identifying as different models by different makers as a basic identity hallucination. They train on enormous volumes of data including lots of people talking about certain makers and models (ChatGPT was actually a super common one given that it became the kleenex of the LLM world). Hence why vendors have to specifically tell it to override that, and if they don't you get lots of funny cases of identity confusion.
This isn't the big gotcha some people seem to think it is, and the whole news cycle about that was mostly by people who have no idea what they're talking about. It's actually a meaningless data point. But it's precisely the sorts of people who think that a few thousand free accounts surreptitiously snuck off with Fable.
cayley_graph 3 days ago [-]
Yup, fair's fair. Anything else stinks of 'rules for thee but not for me' (a maxim the frontier labs seem worryingly happy to apply, on several counts).
ronsor 3 days ago [-]
I am immediately sold on this.
Sorry, OpenAI & Anthropic.
matheusmoreira 3 days ago [-]
> distillation: why exactly is it bad?
Felony contempt of business model.
noncoml 3 days ago [-]
Don’t know much about how distillation works so please enlighten me here.
> what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models
If it’s as easy as that why do they choose to distill another model and not distill the knowledge on the open Internet from scratch?
numpad0 3 days ago [-]
known-good prompt-response pairs are more useful than random semi-coherent texts presumably
paxys 3 days ago [-]
You need to do both.
A model trained on all knowledge from the internet (and other sources) is large but ultimately not very useful by itself, because it is going to spit out all kinds of garbage. You have to apply multiple further stages of training and refinement to the base model before putting it in front of users. So as an example you can train a model by yourself and then have GPT or Claude continuously check its outputs and correct it when it is wrong, ending up with a far more powerful model.
root_axis 3 days ago [-]
Because the model can output data in a manner optimized for training a new model, including outputs that were post-trained like RLHF and RLVR.
qurren 3 days ago [-]
Government cannot exactly "bar" terms of service. ToS isn't law. The most they can do is say they're unwilling to enforce them.
ToS is just conditions that you agree to in order to use a private service that is provided at-will. I can have a private coffee shop where the terms of service are that you must wear red to enter, and if you're not wearing red, you are not welcome on my property.
So it would be upto OpenAI and Anthropic to enforce them on their own terms (by banning accounts and IPs).
ascorbic 3 days ago [-]
The government absolutely can pass laws that ban particular contract previsions. They do that all the time. In your analogy for example while they can require you to wear red, they can't require you to be white.
onesociety2022 3 days ago [-]
Governments can do anything they want by passing a new legislation. In your example, they could easily pass a law that states that any ToS cannot reject service to a customer based on the color of their attire. In the USA, it's obviously already illegal for a business to reject service to a customer based on some protected classes like race.
ButlerianJihad 3 days ago [-]
The joke is on you! I’m not wearing any attire! Hahaha!
nl 3 days ago [-]
That's just not true. You can absolutely have terms of service that are illegal, and the government can enforce them.
magarnicle 3 days ago [-]
Why would reading copyrighted material ever be an issue anyway? Wouldn't copyright law only apply to what you create and publish using the model? Training on every comic book should already be perfectly legal, as long as you accessed them legally, right? But publishing your own Batman comic using that training is copyright infringement.
What I'm saying is, doesn't the law already cover 1?
_aavaa_ 3 days ago [-]
Fair use requires more than you accessing the material legally.
In the US one of the factors is “ the effect of the use upon the potential market for or value of the copyrighted work”.
If anthropic Hoovers up the world’s books and trains on them, and then spits them out verbatim on command, then it will clearly impact the value of the work; nobody will buy the original, they’ll just ask Claude.
Others also argue that even if it’s not reproducing it exactly that the training runs afoul of that factor, specifically the “market for” portion. A rights holder can no longer license their book for training of LLMs if Anthropic goes ahead and just trains on it anyway.
magarnicle 3 days ago [-]
> If anthropic Hoovers up the world’s books and trains on them, and then spits them out verbatim on command, then it will clearly impact the value of the work; nobody will buy the original, they’ll just ask Claude.
Ah, right. So if we want models to be capable we need them to be trained on as much as possible, yet we also want to stop what you described. So what can be done?
_aavaa_ 3 days ago [-]
I mean the choice is: 1) we pass laws that explicitly say training models like this is legal (the original quote, 2) say it’s illegal and requires licenses for the data and ability to opt out, 3) we ignore it and continue because the companies are too big to jail.
petilon 3 days ago [-]
[flagged]
altruios 3 days ago [-]
This is a silly perspective, inaccurate, and out of bounds framing.
Public libraries, in this instance, is curated data from all the internet, obtained through not legal means (I don't have a problem with this other than lack of attribution, being copy-left). Just to be clear.
But in answer to your incredibly leading and inaccurate framing... they are required (by their job title) to teach to those who who show up in the classroom, it's not their place to discriminate against anyone/thing (even those like itself (other robots)) that also show up in the classroom.
But you can't teach at a university using only knowledge learned from the library. you need a degree. You are free to teach at the park, where anyone can hear you. public in -> public out.
petilon 3 days ago [-]
If a professor learns from multiple books, generalizes from them and then shares his knowledge he is providing a valuable service. Versus someone who makes a recording of the professor's lectures and resells them to undercut the professor--that guy is not providing a valuable service.
tux3 3 days ago [-]
What kind of fresh hell does the sentence "undercut the professor" come from?
Teaching isn't a race to the bottom. You don't undercut teaching by giving more lessons, just like you don't slight the hospital by performing CPR.
petilon 3 days ago [-]
We are not really talking about teaching here.
tux3 3 days ago [-]
We sure wouldn't be, if we picked better metaphors.
_aavaa_ 3 days ago [-]
No, we’re talking about an intimate set of tensors, not a human being.
A tree falling and killing someone isn’t tried for manslaughter.
So I don’t care about a hypothetical teacher.
lelanthran 3 days ago [-]
> If a professor learns from multiple books, generalizes from them and then shares his knowledge he is providing a valuable service. Versus someone who makes a recording of the professor's lectures and resells them to undercut the professor--that guy is not providing a valuable service.
I'm confused now; isn't the LLM that trains on that professor's lectures, videos and textbooks undercutting him?
Where were you going with this?
idle_zealot 3 days ago [-]
What about a student attending lectures of other professors and generalizing what he learns from them, then going on to become a professor?
petilon 3 days ago [-]
If the student is really good at generalizing we wouldn't even be having this debate because he would've just generalized from the same source materials the professor used.
altruios 3 days ago [-]
What are you even trying to say: "undercut the professor"...
The further you try to constrain this topic into this illformed analogy the weirder it becomes. If we start off with a better analogy...
Crisco 3 days ago [-]
No, but the students that learn and distill what the professor teaches are not obligated to use that information only how the professor wants them to.
petilon 3 days ago [-]
Can the professor refuse to teach some students, or must he teach all comers?
xboxnolifes 3 days ago [-]
One can teach whoever they want to or don't want to. If one joins a university, that changes things. They are now part of an organization larger than themself.
lostmsu 3 days ago [-]
glhf under the circumstances
smarf 3 days ago [-]
'why is reselling stolen stuff bad'
if the professor took all human knowledge, much of which was explicitly not free, and used it to make a for-profit knowledge machine that extrudes unreliable summaries of that knowledge, then yes, being obligated to teach for free would be a fitting punishment.
mywittyname 3 days ago [-]
More like, is a professor who learned from books prohibited from writing his own books on the subject?
petilon 3 days ago [-]
He is prohibited from regurgitating source material, of course! But if he generalized from the books he read and really learned the subject--and even made new connections between ideas--then he is free to write his own book.
mywittyname 3 days ago [-]
He is not prohibited from "regurgitating source material" in many cases. Facts are free. It doesn't matter who first measured Young's Modulus of aluminum, anyone may state that fact as originally presented.
The professor is free to lift all the facts and formula they want. They just need to rephrase explanations. Which is pretty much what an LLM is going to do.
3 days ago [-]
_aavaa_ 3 days ago [-]
Who cares, a LLM isn’t a person.
janalsncm 3 days ago [-]
1) No one is asking Anthropic to give tokens for free, but at market rates.
2) Any professor who tried to ban students from posting lecture notes online would be immediately mocked.
bluegatty 3 days ago [-]
Making an LLM from raw data is value-add.
Distillation is just value extract.
It's soft, and I'm not sure what the answer should be ... but I think that there is a difference.
I think we start by recognizing that ... and then try to figure it out from there.
'The Internet' may be a public good, maybe we make them pay a tax for that, but that's different than distillation.
scotty79 3 days ago [-]
> Making an LLM from raw data is value-add.
> Distillation is just value extract.
There is a value-add in selecting the valuable parts out of the garbage. And let's face it. Largest models contain a lot of garbage.
bluegatty 3 days ago [-]
I think that's kind of fair, but it still fits within the context of 'some things are value add' and 'more or less than others'.
We ought to identify that and integrate that into our thinking.
nemomarx 3 days ago [-]
What makes the Internet raw data in a different way? wasn't it mostly worked on by people first?
bluegatty 3 days ago [-]
There is value add in AI irrespective of how the data got to what it is.
Literally the biggest thing of our generation - AI - is the living embodiment of that 'value add' writ large.
'What is the difference' - is the AI you use all day, in comparison to 'all the world's data' you can use for stuff and do 'whatever' with it, but are not likely to come up with something hugely useful otherwise. Maybe, not likely, if you did, it would be 'value add'.
nemomarx 3 days ago [-]
Okay, so if the chinese models are used everyday, do they become a value add? Like what's the line you're drawing here. Amount of value it creates?
bluegatty 3 days ago [-]
Designing and creating an LLM from nothing is a monumental feat of Engineering and 'value add'.
Copying something is not.
Programming Microsoft Word is value add, copying the code is not.
Copying design ... there are some question marks there.
It's extremely easy to understand at it's core.
What makes it hard, is that faux intellectuals like to deconstruct ideas at the margins, and have those critiques stand in for reason.
"At sunrise the sun is only 'half there' ... there fore there is no 'day and night' just a blur! Day and night are the same thing!"
The training data used is part of all of this is a separate but related question.
inigyou 2 days ago [-]
they all just copied the Transformers paper anyway
bluegatty 2 days ago [-]
If all it took was 'copying the transformers paper' then why aren't there 1000 SOTA foundries?
inigyou 2 days ago [-]
cause you also need a trillion dollar supercomputer
bluegatty 2 days ago [-]
So - read paper + 'compute' and you can make a trillion dollar company?
OleksandrC 3 days ago [-]
The article makes a point about agent harnesses being sticky (the supposed moat). I have been building my own agent harness for a while, and I can tell with confidence that the harness almost does not matter, the entirety of the AI magic is the model itself. The harness can be almost barebones (like, for example, mini-swe-agent used for benchmarks), and yet the model still does the task just fine.
So from my perspective, it's doubtful that this is the moat. Besides, for example, Claude Code in particular is so buggy (and always has been).
bze12 3 days ago [-]
By the harness I believe he means the entire end-user product experience, not specifically the harness code. I’ve mostly stuck with codex because their Mac app is better and I’ve gotten used to running automations through it. The more workflows they can build around this (design tools, collaboration, etc), the better chance of lock-in.
> If you own the user touchpoint, then you have meaningful lock-in, and the best way to own the user touchpoint is to be the canvas for everything they need to do. This, by extension, means that the frontier labs are on a collision course with software companies: it’s software that owns the user touchpoint, and it’s in the frontier labs’ long-term interest to not simply be a commodity input into software but to simply replace software outright.
Sol- 3 days ago [-]
For me, harnesses are mostly sticky insofar as the model providers only allow you to use their subsidized plans through their own harnesses, unfortunately. But of course switching model + harness is an option.
johanyc 3 days ago [-]
> model providers only allow you to use their subsidized plans through their own harnesses
true for anthropic, not true for openai.
dansquizsoft 3 days ago [-]
Facts, I was able to code a personal self improving harness in a weekend (something a bit more similar to Hermes or OpenClaw at the time but with a more expansive set of features for my use cases and requirements) and it works great for 90% of the tasks I would use Claude Code or Codex (now ChatGPT App) for, with the remaining 10% being able to be implemented with a few more prompts from within the harness itself.
For this reason alone I would also argue that the idea about an agent harness being sticky is a non-starter long-term.
hdz 3 days ago [-]
The harnesses will tend towards commoditization, but for now the harness quality matters a lot. Especially for non terminal harnesses.
colinsane 2 days ago [-]
we've barely scratched the surface when it comes to the available design space of agent harnesses.
i also expect you'll see markedly different results if you constrain yourself to small models. there even trivial harness improvements like Codex's /goal feature, and more capable basic tooling (e.g. semantic code grep, js-capable `fetch` tooling) make or break the actual task success rate.
neutronicus 3 days ago [-]
Yeah it certainly feels like the harnesses are pretty minimal value add on the token pipe
stymaar 2 days ago [-]
> The defining characteristic of a commodity is that it is fungible: a gallon of oil is a gallon of oil; a ton of copper is a ton of copper; a bushel of wheat is a bushel of wheat.
The concept of “commodity” as defined above is a model, a simplified abstract representation of reality, but that does not match the reality perfectly (the map != the territory).
The author claims that a token isn't literally an ideal commodity, but neither is oil or wheat, many factors influence their real value (intrinsic properties, location, available storage at production, expected delivery date, etc.) so that no two gallons of oil in different contracts have the same price.
Is treating “tokens” as a commodity a worse model than treating oil this way? It depends who you ask! I'm pretty sure that a chemist working at a refinery would be more happy to see tokens being felt with like a commodity by his company than if they started viewing crude oil like one.
(Overall, there's way too much economism in that post, and way too few facts, and as a result the argument makes very little sense, the author basically wrote that both OpenAI and Anthropic are drowning in cash right now because compute scarcity means the price must be significantly higher than the marginal cost…)
sanderjd 2 days ago [-]
Yeah I bumped on this as well. Contra the author's claim, the analogy to energy commodities seems very direct to me. Natural gas is not useful in and of itself, what is useful is the energy or aggregates created from it, and those have very different levels of efficiency. Exactly like Sol more efficiently converting tokens into intelligence than Kimi, a combined cycle gas plant converts gas into electricity more efficiently than a simple cycle gas plant. But this does not imply that gas is not a commodity. And both the more efficient and less efficient kinds of plants have large markets; they just target different trade offs.
Edit to add: I think what he's saying is more like "tokens aren't the interesting commodity, 'intelligence' is", which makes more sense. To carry on my gas and electricity analogy, I would say the same thing about gas being the less interesting commodity than electricity, because electricity can be used for a broader set of useful things. But both things are commodities, despite one being an input and the other being an output in this case, and the conversion efficiency is one very important consideration, but not the only one.
miyoji 2 days ago [-]
"Intelligence" isn't a commodity at all. It doesn't make a lick of sense.
The value of intelligence is that it can solve my specific problems in ways that are satisfying to me. The example the article uses is a CRUD app - but the CRUD app I need isn't fungible with the CRUD app you need! It's not fungible at all, it's a specific solution to a specific problem that may have zero value to anyone else, and certainly cannot be replaced with anyone else's solution to their problems.
If we're comparing electricity to gasoline, then models are cars, and "intelligence" (I disagree that this is what LLMs produce, but whatever) is distance traveled.
Distance traveled is not a commodity. It's the desired outcome.
sanderjd 2 days ago [-]
I think he's creating a definition for an abstract thing here, and using the word "intelligence" to represent that abstract thing.
In your analogy, the "CRUD app" is analogous to the car, but that's not what Thompson is defining to be the unit of "intelligence". He's saying that some number of units of "intelligence" are necessary to create that CRUD app, somewhat analogously to how some number of units of energy are necessary to create a car.
But I agree with you that this concept of a "unit of intelligence" that he's using is probably too abstract to ever be usefully well defined.
throwaw12 3 days ago [-]
Lets do "who's afraid of US models" version:
* Me, as an individual, because I might not be able to pay price hikes, because my revenue (salary) is much lower than what they want and I can't support my expenses via huge bank loans.
* Again, me as a new entrant to the industry, LLMs are basically pay-to-play games, again related to price hikes, new entrants might not be able to afford paying those prices 24/7 - which you need when learning new things.
* Any non-US company, US can block the models which can disrupt the whole business.
* Even some US companies, for example if you operate in EU and EU somewhat changes their mind and follow the ICC and require you to stop working with Netanyahu (war criminal as per ICC), then following laws in EU, might create trouble to your whole business.
1over137 2 days ago [-]
Also me, as someone who lives in Greenland, Canada, Venezuela, Cuba, Iran, etc. China is not threatening to invade, USA is.
rektomatic 2 days ago [-]
you live in all those places??
ryoshoe 2 days ago [-]
By this same logic there's nothing to be scared of in regards to Chinese models if you don't live in Taiwan
jbstack 2 days ago [-]
You might take a different view if you live in Taiwan.
traceroute66 2 days ago [-]
> Any non-US company, US can block the models which can disrupt the whole business.
And read your data, see CLOUD act, PATRIOT act etc. etc.
No longer a theoretical risk in today's US political environment.
mc32 2 days ago [-]
Hasn’t that been the assumption since Carnivore? Also we tapped European pols phones back in ‘12, was it? Though it’s not like the Eu are a bunch of innocents either…
It’s not a new thing. Industrial espionage has always been a thing as well. So has bribery (for deals) been a thing especially by euro concerns.
traceroute66 2 days ago [-]
> It’s not a new thing. Industrial espionage has always been a thing as well.
Well sure, except with closed US LLMs you're basically just handing them data on a plate, and paying for the privilege. ;)
Very American really ... monetising industrial espionage.
2 days ago [-]
radicalbyte 2 days ago [-]
Our intelligence services have capabilities, it's not just the US. The Patriot / Cloud Act increase the scope outside of intelligence agencies. I'm of the opinion that for military intelligence there is a legitimate need to have some - black - capability. It should just be highly limited in scope.
traceroute66 2 days ago [-]
> It should just be highly limited in scope
"should" is doing a lot of heavy lifting there.
I agree it is hard to escape some form of legitimate need for intel.
The concern with intel and the present US administration comes on the checks, balances and controls side.
We are after all dealing with an administration happy to conduct much of its most sensitive business on Signal using off-the-shelf phones.
mc32 2 days ago [-]
I don’t think these issues only exist or have existed in this admin. Previous admins were tapping the phones of the heads of state of Europe. There was also mishandling of intelligence. Perhaps the magnitude is changed though I don’t have evidence for that.
Some of it I think is selective picking. Similar to election interference where we know quite a few foreign states like interfering with our elections but we typically only hear about one of those states as being the culprit. Now, obviously we like interfering as well (Ukraine in ‘14 and Iran today, though Iran would be less controversial) and many others over the years.
sneak 2 days ago [-]
No. Lots of people have never heard of FAA702 (which has been well-documented in mass media) much less Carnivore or ECHELON.
The default email provider for most people in the west is Gmail.
2 days ago [-]
amiraliakbari 2 days ago [-]
Also me, as someone who is living in a place that US may drop bombs on because the models may think it is a military target, or even a higher priority target like a girls school.
ninjagoo 2 days ago [-]
> Lets do "who's afraid of US models" version
Ah, cultural nuances. The title "Who's Afraid Of Chinese Models" is a riff on "Who's Afraid Of Virginia Woolf" which itself is a play on the song "Who's Afraid Of The Big Bad Wolf".
The title essentially means that the chinese models are being portrayed as the big bad wolf; but are they really the threat or are american frontier labs afraid of competition and commoditization?
It's also somewhat ironic because the author says that there is something to be feared -that western innovation will become dependent on chinese models, especially for cyber, if the american ones are restricted or unavailable.
Kuyawa 1 days ago [-]
Me as in "Unfortunately, Claude is only available in certain regions right now"
DeepSeek, Kimi, Xiaomi Mimo, Qwen, Minimax, GLM, Hy3 and Ernie are always available, and I can't be happier
therealpygon 2 days ago [-]
I think your second point touches on a bigger concern for small businesses.
US labs have consistently demonstrated their intention to paywall higher intelligence. Eventually the paywall for “hyper-intelligence” will set a bar so high that average small businesses simply won’t be able to afford the bill of what is used by the top corporations to keep themselves at the top. That’s already starting, when it comes to the volumes of tokens top corporations are burning.
This is a feature of the system corporations want to establish and OAI/Anthropic are happy to oblige. 10% of a trillion dollar company is the same as 10% of 1,000,000 million dollar small businesses. Whose hitch would they rather ride, and which size customer easier to obtain to meet their revenue goal?
Not to mention, it would certainly be possible for the EU and other world powers to equally disincentivize use of US labs as a data security risk, since our top models are impossible to run in private lab environments without specialized agreements and there no access to the model weights for auditing. As best as I can tell, AI regulation is a dangerous game that is a hair away from isolationism.
In my opinion, Google is one of the few hopes in this area. There is still a paywall, but I feel like they are the closest thing to a Chinese lab we have (for frontier) in terms of their targets (real business use cases) and they actually have both the infra and already have a pipeline for small businesses into their products; they already have the wide non-AI customer base to leverage unlike Anthropic and OpenAI whose only product requires convincing people to use their (more expensive) AI.
kiicia 3 days ago [-]
in simple terms closed models are rugpull waiting to spring at unsuspecting users, it's like you take all worst components of terminal capitalism (including price cartels), subscriptions relying on almost monopolistic dependency and microsoft/uber models of hugging competition to death to remain only provider and dictate all conditions
aa-jv 2 days ago [-]
For me, the most important factor is:
* All modern AI is a perfect front for harvesting material for processing by NSA/GCHQ.
Given the criminal US' 5-eyes/9-eyes apparatus' atrocious war crimes and human rights records, this is reason enough to eschew American AI 'products'.
I'll use the AI created by the culture that lifts a billion people out of poverty first, not that from the culture that murders children every 15 minutes and lies to itself about it ..
girvo 2 days ago [-]
As someone from neither the US nor China, neither are altruistic and both have horrible histories, both meddle in my country’s affairs, and neither have my interests at heart.
China’s got plenty of blood on its hands, pretending otherwise is silly. America does, too. It’s quite easy for me to condemn both their governments and trust neither.
aa-jv 2 days ago [-]
The USA has far more innocent blood on its hands, this century, than any other nation state.
And only nationalist/racists ignore the fact that China has lifted a billion people out of poverty during the same period that the USA and its allies have destroyed countless other sovereign nations and left 36 million refugees from their illegal wars for the world to deal with...
China has a lot, lot better record on human rights than any member state of the criminal 5-eyes/9-eyes countries, which engage in massive human rights violations at atrocious scales every minute of the day.
girvo 14 hours ago [-]
If you want to talk about this century and human rights, then you better look inwards. You're obviously a nationalist yourself, accusation in a mirror as is typical.
tristanj 3 days ago [-]
The people who are most afraid of Chinese models are the VCs who poured into Anthropic and OpenAI at astronomically high valuations. Anthropic is valued at $1.2T and OpenAI is targeting $850B. These astronomical valuations were built on the premise that these labs would generate massive profits from premium API pricing, but the Chinese labs are completely undercutting this strategy by releasing excellent open models for free. If the frontier labs are forced to cut prices and join the race to the bottom in token prices, these valuations are unjustified, and VCs will face enormous (paper) losses.
mediaman 3 days ago [-]
The (quite excellent) article discusses several of your points. If you haven't read it, I recommend it.
- Commodity market profitability is determined by marginal cost of production. LLMs have marginal cost; traditional software does not.
- Models are not free. Downloading them is free. Running them is not. This has manufacturing economics, not software economics; the idea that they are "free" is an economic category error as it relates to their actual use
- The highest tier Chinese models are not more economical than US frontier models. Try GLM 5.2 and see how much it costs to do real work. I did, and it was more expensive than GPT 5.6.
- This is because US labs are leading on cost efficacy of inference ($/task)
- Training will decline as a percentage of costs as inference expands compute share due to agentic workloads. A big part of training now is optimizing token efficiency. It's hard to distill token efficiency; that is perhaps why Chinese LLMs are so inefficient.
- With increasing inference as % of total compute, if labs create efficient models -- which they can, because they can create highly optimized models amortized over very high inference loads -- they can be low cost producers, and be competitive at $/task rates
OpenAI really shows the way here. Their cost per task is less than half that of Anthropic because of more efficient tokenization and less verbosity. OpenAI is both cheaper and better than Chinese models for frontier work.
larrysalibra 3 days ago [-]
Ben's article "distills" down to 2 reasons that US frontier labs shouldn't be "afraid":
1. US frontier lab unit economics are better
2. US frontier labs are moving up the stack making tools that are "stickiness" and will prevent users from switching.
For 1...he doesn't provide any evidence for US lab unit economics being better...the major input to unit economics is electricity...which is cheaper in China. And building data centers and connecting them to electricity is both cheaper and an order of magnitude faster in China. The main input that US labs might have an advantage in is in cost/access to chips, but that given the level of chip investment in China it seems unlikely to hold.
For 2...there's little evidence these tools are sticky. At least in programming, the trend seems to be tools like opencode that support multiple models and providers.
And even when they are sort of sticky, as we know on hacker news, people figure out how to point the tools they like to competing models even when the app doesn't official support it.
And every improvement in model capability makes it increasingly easier to make your own tools.
He’s glossing over the reason they are not: 90% profit margin of Nvidia. Power is only a small part, single digit, it will eventually matter but does not really today.
What is the cost of AI? The single largest ingredient is Nvidia profit margin.
Huawei accelerators are not as efficiency yet, but they don’t nearly extract as much margin.
Why would future revenue stay with the labs given this situation? This whole thing had an airline industry sized red flag on it that makes investing into frontier lab about as sexy as investing in United.
Maybe the token economy is some kind of reverberation of the airline reward miles economy, the emergency hatch to be able to survive under maximal supplier extraction (Nvidia is just the top of a monopoly stack here, even if they replace those chips, the HBM, ASML, Foundry layer can get their dues)
michaelt 3 days ago [-]
> Power is only a small part, single digit, it will eventually matter but does not really today.
Sorta yes, sorta no.
A single 5090 consumes 450W - at Californian energy prices of $0.38 per kWh that's $0.17 per hour. And the card itself costs $4100 on amazon. So after 2.75 years running at full power 24/7 you'll have spent more on electricity than on the card. I would have thought most data centres being built today would have a design life longer than 3 years.
Of course you can throttle the cards to ~300W without losing too much performance. But also you need more than a single 24GB card to run most modern LLMs.
So about 300x the 5090's power but 1000x the price. Roughly 9 years for electricity to exceed price at $0.38 and datacenters will show up in areas with cheaper power than CA.
Omniusaspirer 3 days ago [-]
Anyone seriously building out AI infrastructure I presume is paying nowhere near $.38/kWh which is extortionate. Utility scale solar is closer to $.02-.03/kWh, then maybe around ~$.10/kWh for natural gas peaker plants.
londons_explore 3 days ago [-]
Utility scale power price varies widely by location and exact time of day.
LLM's aren't very latency sensitive and can therefore move to wherever power is cheapest.
Right now that's places next to aluminium smelters (which also like very cheap electricity 90+% of the time).
inigyou 2 days ago [-]
GPU design life is commonly cited as about 3 years. The main driver is not physical degradation of the cards but the expectation they'll become obsolete with newer cards.
threatripper 3 days ago [-]
Also add electricity 50% on top for cooling the DC.
aurareturn 3 days ago [-]
Cost of electricity isn’t a long term advantage in my opinion. Private companies will figure it out.
What matters most is $/completed task. It does seem like OpenAI and Anthropic are winning here even with worse electricity rates. Perhaps it is made up by the efficiency of Nvidia and Broadcom chips, which China can’t get in mass.
I do think that OpenAI and Anthropic are moving up in stickiness. My company has rallied around Claude. We are customizing Claude Code, adding knowledge bases for non technical people, writing skills for them, using Claude features company wide. It’s hard to move.
Meanwhile, I personally use ChatGPT outside of work. The memory, ease of use, habit keeps my subscribed.
blensor 3 days ago [-]
I'm a model nomad, using whatever solved my last problem the best and where it makes the most sense to start my next work in.
However with the latest models Fable, Kimi K3, 5.6, it's getting to a point where I sometimes forget what model I am on without noticing a difference. And once I realize it because something may not be exactly like I expected it I won't switch for that work either because I don't want to invalidate the cache.
For the next work I will do there is maybe a 50/50 chance to remember to switch the model before I start.
That's not what I would call stickiness towards a certain provider.
amelius 3 days ago [-]
You didn't say what kind of problems you solve with AI. It matters a lot if you are doing HTML versus C++, for example.
blensor 3 days ago [-]
In no order of importance:
- Refactoring a 13 year old in-house vacation rental booking system ( python/turbogears )
- Backend development for our VR fitness game ( flask/python )
- Unity development on our VR fitness game ( C#/Unity )
- VR game development experiments ( Godot/GDScript )
- Standalone SLAM localization service ( C++ )
- Audio analysis ( python/pytorch )
- Virtual display with Viture display glasses ( C )
- Reverse engineering a library I am using for another project ( ghidra -> C - no MCP yet, that's something I am looking forward to )
- Public facing website rebuilding for the booking system above ( PHP/JS )
- Generative 3D environments for our VR fitness game ( python )
- Wireless camera/IMU based tracker for the SLAM system ( C )
Once I've dug in with a specific model into a problem I tend to stick to that because I have a feeling what it will do and how well it works, but when I start a new thing I usually use whatever the model was last set to.
amelius 3 days ago [-]
Wow, now we're talking :)
wbadart 3 days ago [-]
Seems like most popular harnesses, including codex and Claude code, support Agent Skills (an open spec for skill formatting/ organization): https://agentskills.io/clients
Which is to say, this isn't really a lock-in/ stickiness vector (unless maybe the wording itself of a skill is hyper-optimized for a specific model)
littlecranky67 3 days ago [-]
you can simply switch to z.ai/GLM-5.2 inside Claude Code by settings env variables in .claude/settings.json
culi 3 days ago [-]
> Private companies will figure it out.
Across sectors, China added 543 GW of energy in 2025. Next year, USA is expected to add between 70 and 80 GW of energy
metalspot 3 days ago [-]
You really have to look at energy/capita and how much energy is embedded in exports. The gross numbers are misleading. The US wasn't building new electrical generation capacity because it didn't need it and there was no market for it (caveats apply, but in a broad sense this is the major reason). Now that the market exists the question is how much can the US actually bring online and how rapidly, which is a real challenge after decades of degrowth politics used to justify slash and burn consumption of the industrial base.
AI is really all about electricity. AI could be completely fake and yield zero value whatsoever and the US would do exactly what it is doing now because the AI bubble is what creates the market for building new electrical generation capacity, which is needed for re-industrialization. Also why our friends in UK/Europe/China are so busy pushing anti-AI propaganda to try to undermine this.
culi 2 days ago [-]
> The US wasn't building new electrical generation capacity because it didn't need it
You need to understand that the US added 40.3 GW of energy capacity in 2023. 70 to 80 represents a dramatic growth (nearly doubling) in new capacity.
If AI is "all about electricity" as you say then the US was barely ever in the running in the first place. There's no possible way the US could compete any time in the next 2 decades (putting aside a dramatic shakeup like war)
> Also why our friends in UK/Europe/China are so busy pushing anti-AI propaganda to try to undermine this.
Lol people make up the funniest theories when a political idea they don't like is gaining popularity. At least you didn't blame Russian bots
21asdffdsa12 3 days ago [-]
Oh, the UK and Europe castrated themselves and want others to follow the example. They still don't have made the connection between "I give up ability" and i get attacked by a proxxy opponent by those i gave ability up too.
They do not want to life in the world that is and thats going to be, but in the past and the world they green ideology promised. Reality denial be a addictive poison.
flir 3 days ago [-]
That last para is a novel idea. I wonder if there's any evidence to support/undermine it though?
metalspot 2 days ago [-]
The concepts of industrial reserve capacity and using dual-use consumer goods to subsidize military production capacity are well known and widely practiced historically in the US. China adopted this strategy from the US, and the US conveniently forgot about it for a few decades in order to justify selling off the industrial base to China, but at least based on public documents like the published U.S. National Security Strategy, I would assess with high probability that this is explicitly recognized and being followed now.
AI is not fake and it does work, but what I am saying is that from a pure systemic analysis perspective, you can do the numbers, and even if AI was complete fugazi, the benefits you get from the electrical generation capacity, and the ability to fund it through private markets, which bypasses Congress, and locks in commercial contracts (often with foreign governments) which will be almost impossible politically to reverse, would still make it optimal from a strategic perspective. That is my calculation, and to the extent that it is correct, I would assume that the US Military's strategic planning apparatus would arrive at the same conclusion.
AI compute has some unique characteristics that make it especially useful for grid management. Moving consumer compute to the cloud means that the electrical use of that compute can be centrally managed. In an emergency, you can cut electrical use for consumer AI by 50% or more, because chips run more efficiently at lower power, and you can shift workloads onto quantized models, reduce resolution for video output, etc, to reduce compute, which leads to minor service degradation but not interruption. AI datacenters are also adding massive amounts of battery storage capacity, which is an additional grid buffer. For every GW in capacity added by hyperscalers that is creating a dispatchable reserve capacity of 50% under completely normal circumstances (hyperscalers do this internally to optimize their own costs) and then that number goes up depending on the scale and duration of the emergency.
anonzzzies 3 days ago [-]
They are not stopping either.
golem14 3 days ago [-]
I'd really love to see the evidence on this!
HarHarVeryFunny 2 days ago [-]
> 1. US frontier lab unit economics are better
That's not generally true, since there is generally still much reliance on NVIDIA. The true low cost providers are Google with their TPU and vertically optimized stack, and Amazon with Trainium. However, Google does not have their own frontier model, and Anthropic (who are partially served by Amazon) are also paying a premium for extra NVIDIA-based capacity from SpaceX, maybe soon from Meta too.
I don't know how the economics of domestic Chinese Huawei-based clouds (no NVIDIA) compares to the west, but since serving cost is mostly hardware depreciation and to a lesser extent electricity, they are not necessarily at a disadvantage (Ascend 950 costs roughly 50% of an NVIDIA H100), and more to the point it is irrelevant when considering US commercial use that is more likely to be using Chinese open weights models from US providers served on NVIDIA based hardware.
I think the real significance of Chinese frontier models being open weight is that it takes development cost amortization out of the US-based serving cost, while the US AI labs can't afford to do this. The US labs therefore need to reduce development spending to remain price competitive. The Chinese companies are of course still making money from the Chinese market, whether by selling API access or by other business models such as Ziphu making 75% of it's total revenue by selling services to Chinese customers who are running their models on-prem due to the Chinese apparently being very concerned about data privacy.
pishpash 3 days ago [-]
More basically, production cost matters only if inference is priced at commodity prices. That's not what VC's signed up for, which is rent-seeking.
Computer0 3 days ago [-]
In a corporate setting yes Opencode all the way. However in a non corporate setting I am getting $3000 of api usage a month for $100 at Anthropic and only use open code for the smallest cheapest tasks
tvink 3 days ago [-]
I think calling opencode the trend is naive. This not what is being run on company time.
hack1312 3 days ago [-]
OpenCode is absolutely used on company time.
sciencejerk 3 days ago [-]
Shhhhhh...! OpenCode only runs on authorized machines by responsible employees following company policy ;)
striking 3 days ago [-]
Sure, let's have a look...
> I highly doubt that Chinese models are cheaper to serve on a marginal cost basis, they just seem cheaper because Anthropic and OpenAI are so supply constrained that they are charging far more than they would if there were sufficient supply to meet the demand for intelligence. [emphasis mine]
I guess I'm missing the part of this article where they bring hard numbers in to back up the argument here. What work was attempted? https://cursor.com/evals shows the previous generation of open models (Kimi K2.7) trading blows with the others, cost effectively. Composer 2.5 is itself a fine-tune of K2.7, and it's apparently quite token efficient, so why would it be impossible for a Chinese lab to achieve something similar? GLM 5.2 Max is also ranked above the lower end OpenAI models and is not far off in price.
It's weird to have this entire discussion about tokenomics without mention of the circular financing and debt raised by labs in the West, which can then essentially give away their capacity to end users. OpenAI giving away quota resets to subscribers like candy on Halloween while their compute partner Oracle's bonds is reevaluated to be one grade above junk? How?
I don't think you can make an argument about the future one way or another by arguing using the listed prices. The math is not internally consistent enough for it.
gruez 3 days ago [-]
>What work was attempted? https://cursor.com/evals shows the previous generation of open models (Kimi K2.7) trading blows with the others, cost effectively
Because you're comparing retail price whereas the parent commenter (and the article) is talking about marginal (ie. inference) costs. American labs are providing a premium product and they're charging accordingly. Meanwhile for chinese models they're open weight so they're limited to how much they can charge without competitors undercutting them.
If we use tokens as a rough proxy of inference costs (rough approximation, I know) and look at artifical analysis benchmarks, you see that all the open models are behind the pareto frontier in terms of efficiency.
striking 3 days ago [-]
I'm arguing we can't trust retail prices because the marginal pricing isn't meaningfully connected to it anyway.
But if we have to look at what we think margins might look like, DeepSeek continues to host v4 Flash at the existing price despite competitors beating it in price (https://openrouter.ai/deepseek/deepseek-v4-flash), so there's at least one example of a Chinese lab charging a predetermined price despite competition. And no one but Moonshot is hosting Kimi K3 yet (https://openrouter.ai/moonshotai/kimi-k3). Perhaps there's room in the market for those who release their models to make margin on them.
And I believe my Composer example speaks for itself. The open models are behind but there's tangible proof they can be tuned for pareto frontier efficiency. See "Cost per Task" at https://artificialanalysis.ai/agents/coding-agents.
their competitors are discounted at around 33%, so it's safe to say that's the margin, maybe less if their competitors have worse caching or quantization. Meanwhile claude code/codex resellers selling tokens for 90% off API price, presumably by reselling usage from fixed consumption plans, which gives an idea on how fat the american labs' margins are.
>And I believe my Composer example speaks for itself. The open models are behind but there's tangible proof they can be tuned for pareto frontier efficiency. See "Cost per Task" at https://artificialanalysis.ai/agents/coding-agents.
But composer is a closed model? If it's really that easy to get better coding performance, why haven't the chinese labs replicated it? And this is all assuming the performance boost is real and not from benchmaxxing. Moreover if you apply the "street price" discount I mentioned above, American labs look far more favorable.
striking 3 days ago [-]
The fixed consumption plans are offering several times their worth compared to API pricing with completely free cache reads: https://she-llac.com/claude-limits
I look at that and think that they must be losing money hand over fist on something like this, not that this shows what their margins are like. If their margins are like this then I don't see why they'd be raising money and shuffling it around in circles.
> If it's really that easy to get better coding performance, why haven't the chinese labs replicated it?
Nobody said it would be easy! I just think it's possible, and that presumably they will get around to doing it at some point.
ycui7 3 days ago [-]
discounted competitor could cheat. they can offer subpar model response and sell it as deepseek-v4. it is uneconomical to prove inference providers are cheating, so they get away with it. cheating inference provider does not care if their customers stay.
c0brac0bra 3 days ago [-]
The 33% discounted competitors have no non-retention policy
byzantinegene 3 days ago [-]
the reselling of fixed price plans by resellers are causing american labs to lose alot of money and that's why they're trying really hard to stamp it down.
metalspot 3 days ago [-]
Deepseek is still charging cached input at 1/10th the price of any competitor.
For an example of my real token usage for a day with DS: Input (Cache hit) 530,949,760, Input (Cache miss) 7,875,004, Output 1,389,685 - it is still 1/5th the price of Baidu (the cheapest) and 1/7th-1/10th the price of US hosts.
Also, Deepseek platform is not the same thing as Deepseek open weights. There is a major misconception that the existence of an open weights model means that it is the same thing as the proprietary platform offering, but that is definitely not the case.
ehnto 3 days ago [-]
US running costs are higher than in China, because the US lags behind in energy, has higher real estate costs, and wage costs are higher.
Eventually we will hit a "good enough for cheap enough" and frontier models will hit diminishing returns (if they haven't already for a lot of types of work)
Don't think the rest of the world will sit on their hands while the US soaks up chips either, demand gets filled and if the US won't fill global demand for chips that's an opportunity to undercut again.
The other thing the rest of the world doesn't have to fund is the ridiculous valuations on these companies.
Unless you think the US can stay ahead just with model efficiencies, and that no one else will eventually match them, you are looking at the writing on the wall.
All that to say, the rest of the world is more than willing to eat your lunch, they have a dozen good reasons to, and they're already showing good results.
Just on the economics side, we've been here before too, US companies typically export their commoditization and live on brand royalties. Think all the cheap manufactured goods, the US doesn't make any of it. That's because the US can't compete on margins for numerous reasons, it's too expensive, I don't think AI is any different here except that the brands are currently valued in the trillions and I suspect that greed will be their undoing.
analyte123 3 days ago [-]
The US does not lag behind in energy. Industrial electricity prices in most places in the US are competitive with China, or even cheaper.
ehnto 3 days ago [-]
Apologies, I meant ability to deploy new generation. You're right that areas of the US are cost competitive.
This can change quickly though, so it's not that big of a deal. If AI energy demands push the Gov to deregulate/fast track new plants, or the industry decides to build out their own generation renewables.
zx8080 3 days ago [-]
Links please.
analyte123 3 days ago [-]
IEA’s chart for “final electricity price for large industrial customers in energy-intensive industries” [1] shows the US cheaper than China, but this is using data from Texas as representative. I’m not sure how they back this representation up, or how different this is from GPP “business” prices. So “most places in the US cheaper than China” may be wrong, but certain states or regions are at least competitive.
This is very misleading and almost certainly deliberately so.
Texas is the only is State that doesn't participate in the national grid (90% anyhow). Their prices are lower because they dont abide by the regulations that they would be required to if they did. Like weather proofing infrastructure.
This is also why their grid crumbles every time it gets cold, or hot, or... Tuesday.
See the 2021 grid failure for a particularly bad example. In that instance somewhere between 250-700 Texans died to keep their prices low.[0]
This seems to show that even with China taxing industrial electricity to subsidize household electric bills it's still 30% cheaper for industrial electricity in China. [0] Those same taxes make household electricity about half the price of US power.
If I had to wager why, I'd say it's due to embracing solar on massive scales recently. Only a few years ago the US was competitively priced.
But the clean, beautiful, coal will come online soon right?
appplication 3 days ago [-]
I think there are some really interesting thought there, but I’d challenge some of this:
> Models are not free. Downloading them is free. Running them is not. This has manufacturing economics, not software economics; the idea that they are "free" is an economic category error as it relates to their actual use
I think a large part of manufacturing economics is illiquid overhead and the cost of expertise to set up and run your manufacturing line. Compute economics don’t have the same illiquidity nor do they require the same expertise or even specialized infra (current temporary chip shortage aside).
The implications of this are small players (e.g. your uncle running an inference server out of his garage) have comparably efficient marginal costs as big players. Compare this to actual manufacturing where small players have essentially no access to the manufacturing facilities of the big players.
Additionally, big players with a lot of compute who are not meaningfully in inference today (e.g. Amazon) have a fairly straightforward glide path to utilizing that compute to compete.
> This is because US labs are leading on cost efficacy of inference ($/task)
It’s possible, but I would need to see better data on this.
>A big part of training now is optimizing token efficiency. It's hard to distill token efficiency; that is perhaps why Chinese LLMs are so inefficient.
I think it’s fair to assume this is true, but also token efficiency is not a meaningful competitive moat. It’s not like these are secrets the Chinese will never figure out, it’s a fairly active research space and the outcomes are quantifiable.
senderista 3 days ago [-]
Amazon is not "meaningfully in inference"? Bedrock seems to have a ton of enterprise customers, some of which would never trust the AI labs themselves with their data but will trust Amazon.
cdud3 3 days ago [-]
A lot of this companies are also in parallel evaluating running Chinese models inhouse to be independent of the good will and profit margins of externals.
This is where Oracles "datacenters for rent to run your own Chinese models" strategy will benefit. The LLM SaaS game is a lost one thanks to China.
appplication 3 days ago [-]
Relative to their other compute or other players inference, not as significantly. Though yes, they certainly have some share.
lemax 3 days ago [-]
But this assumes Chinese models will not achieve token cost optimization. Intelligence needs are fairly flat for many tasks, and the Chinese models have caught up on this front. Next they achieve greater token cost efficiency and we don’t need OpenAI.
overfeed 3 days ago [-]
That the author doesn't acknowledge the relentless R&D efforts DeepSeek has been plowing into optimization, and giving a default win to OpenAI/Anthropic on the supposition that they've been serving models for longer is a black mark against the article.
I appreciate the transparency in explicitly stating their motivation for writing the article (a response to what the author saw as an overreaction to Chinese models), but I feel the article goes too far the other direction, with multiple unsupported leaps of logic, and overstating the stickiness of AI client products.
VulgarExigency 3 days ago [-]
The model that is most optimized around token cost is, in fact, Chinese. DeepSeek is astoundingly cheap by default, but if you use it from Reasonix (the harness optimized around its cache), it becomes even cheaper.
ryeguy 3 days ago [-]
I keep seeing mention of the cache, what's special about it? All frontier llms have prefix caching, what is special about deepseek's approach?
This also comes with significant capability reduction. deepseek-v4-flash is very good in the < 250K range, then degrades between 250-500K, and is practically unusable after 500K.
[edit]
This is my observation from using it without an specific context engineering to optimize for Deepseek's cache compression and sparse attention mechanisms. I am pretty sure that if you specifically structure your context to align to the cache compression boundaries you can significantly improve performance in the full 1M context, but there is not much reason to do this, because if you design your outer loop to work with shorter contexts that solution is portable and more efficient, so I haven't bothered with a optimizing for DS at this point.
Der_Einzige 3 days ago [-]
Btw - assuming NeurIPS reviews aren’t garbage tomorrow, I’ll have a paper out which claims that most long context problems in models are really sampling problems in disguise
Switch to a modern sampler like min_p or ideally a better one like top-n-sigma (it’s in llamacpp) and your “my model gets stupid at long context problems” will basically go away.
Unfortunately this fact is still not well appreciated yet despite nearly every modern sampling technique getting an oral wherever they get presented. Min-K just got an oral at ACL 2026, for a hyper recent example of this. There’s a reason they keep getting orals.
The field massively ignored sampling for mostly safety reasons and now the whole field incorrectly believes long context doesn’t work on small models. Long context is an out-of-distribution problem. Your sampler configured properly keeps you in distribution.
Oh and this is doubly true for quantized models. I run my qwen 3.6 27b with 4bit quants from unsloth and get excellent performance because my sampler stack is good and not the garbage that is top_p and top_k. Also, yes, you need to ignore the trash recommended sampler settings from the Chinese labs (they’re wrong/bad).
metalspot 2 days ago [-]
Interesting. I am not familiar with model internals at this level because I have only been working at the application layer so far, but will definitely research this further. When you get the paper published would appreciate if you can drop a comment with the link so I can read it.
SoftTalker 3 days ago [-]
> Models are not free. Downloading them is free. Running them is not.
Is this really different from traditional software? Downloading postgres is free. Running it is not. You either buy hardware and assume the costs of owning and running that, or you pay to run it in the cloud.
aurareturn 3 days ago [-]
I think the point here is that it takes the same hardware to inference an open source model as OpenAI/Anthropic inference their models.
IE, a lower param OpenAI/Anthropic model can compete with a higher param open source model.
So even if you are an American company who downloaded Chinese models in hopes of saving in cost, you still have to beat OpenAI and Anthropic in $/task which is very tough to do over the long run.
Gigachad 3 days ago [-]
The problem is these AI companies are running at a loss with the hopes of pumping the prices after everyone is hooked. Now they are locked in at selling API access at rock bottom price. The valuations won't hold.
elictronic 3 days ago [-]
It’s even worse. There operating costs are rising because of their own demand claims while those same rock bottom api prices are locked in.
There being squeezed by their own stock pumping and SpaceX pretending to be an AI company is driving down the exit strategy. I’m guessing one starts going full Theranos and begins claiming full AGI or gets the US government to government cheese then hard. It’s going to be a few crazy months.
pishpash 3 days ago [-]
No they don't. Are OpenAI and Anthropic charging at cost, or wish to? No.
aurareturn 3 days ago [-]
If you have more efficient models, you can reduce price and grab market share or keep price and increase your margins.
Over time, this compounds. More profits means more investments. More market share means more control.
charlieflowers 2 days ago [-]
But inference costs scale per task whereas platform services like postgres typically amortize across tasks. If you are selling tasks done by inference, then compute is part of your COGS and it scales per task.
SoftTalker 2 days ago [-]
Much more so, I would agree. A bit like how mainframe batch jobs might have once been charged back based on "cycles" used for a job or other tracked metrics.
ChaitanyaSai 3 days ago [-]
The thing I do not understand here because it seems obvious: AI will be a commodity market and you simply cannot have a large PE multiple. So the valuations imagine a global commodity monopoly or duopoly coupled with the increased intelligence still disallowing other suppliers from becoming competitive? Without any network effects to help?
horacemorace 3 days ago [-]
Perhaps to moneymen the difference between “ChatGPT” and the technology behind it isnt’t obvious. I’ve been very surprised at how few otherwise smart people are completely in the dark about how capable current models are.
As soon as manufacturing starts building this stuff more, it will commoditize. The hardware prices won’t be terribly larger than the original. We’ll have a “Bambu labs” style company to make the AI OS, whatever that is.
jppope 2 days ago [-]
There are many ways to create "sticky" products, even commodity products (e.g. coca cola, starbucks, etc). Network effects are just one.
Regarding Price to Earnings... I'm not sure everyone fully fleshed out the end game for the frontier companies. It appears the pitch is that this current phase is a stepping stone to Artificial General Intelligence or Super Intelligence. I can understand the perspective of investors though... If you can get half of the world on these products at some point you will find something you can sell them even if its not the core product (i.e. loss leader).
bg24 3 days ago [-]
AI (llm) will be a commodity market => I am not sure it was obvious. As of last month, folks thought open weight models are lagging by 6+ months. Once K3 is taken for a deep run across many use cases, it will be clear where it stands. But yes, I agree that now that intelligence is commodity, everything changes.
abernard1 3 days ago [-]
" - The highest tier Chinese models are not more economical than US frontier models. Try GLM 5.2 and see how much it costs to do real work. I did, and it was more expensive than GPT 5.6."
This is a flatly false statement for most things powering backend applications. The AI consumer "doing real work" model, either for analysis, chat, or coding could well be more cost effective with closed frontier models.
But most of these internal glue business SaaS applications where engineers are integrating are not those tasks. It is those tasks which 1) drive immense amount of domain-specific data into the platform over time, and 2) are most encouraging of driving open model independence with no vendor lock-in.
Anyone on this site who has actually used ML models (more accurate in many cases) knows there's a lot of kludge that simply does not need a 5 minute agentic feedback loop to solve the problem. And they were solvable a year ago with lower class models. The token economics are exceptional and the anecdotes of a16z saying 80% of startups are productionizing open models is only surprising to people who think running your company on OracleDB in 2026 is a sound engineering decision.
mediaman 3 days ago [-]
You're correct, but that's a different market segment and not the market GLM 5.2 and its peers compete in.
The labs are not interested in the small, fast, single purpose end of the market. Google increased their pricing on Flash so much that it stopped becoming a cheap model; instead, they released Gemma 4 open source, which is actually easier to use from a third-party inference provider than from Google.
From a total token volume perspective, these "utility" models (classifiers, simple summarizers, small OCR models) will absolutely drive enormous volumes of tokens, at low prices and margin and modest overall market size. Because the models are small and the performance requirements are modest, and because their use cases are specialized rather than general, there are poor economies of scale: they can run cost effectively on rented small GPUs, and a big player doesn't get a structural cost advantage. These models are usually 1b - 30b in size, and can run on a rented 5090. I've productized these myself: I run millions of pages through a fine tuned 1b OCR language model that runs on 5090s at a cost far lower than commercial providers.
But that's not the segment of the market where GLM 5.2, Kimi 3, etc., play. They compete with frontier capabilities, and they are not particularly cheaper than OpenAI models at a cost per task. (I do actually think they compete well with Anthropic, because Anthropic's model efficiencies are poor compared to OpenAI.) And although this part of the market may not be the bulk of the token volume, it is the bulk of the market value.
That's because a lot of human knowledge work is too generalized and fuzzy for dedicated, fine-tuned models, so they are almost entirely different markets that don't particularly compete with each other. (Though if SaaS companies successfully build around verticals that can use small models applied against well-defined jobs, there may be opportunity to push the small/big capability boundary to subsume marginally more valuable tasks that today would require mid-grade reasoning.)
ignoramous 3 days ago [-]
> "The highest tier Chinese models are not more economical than US frontier models. Try GLM 5.2 and see how much it costs to do real work. I did, and it was more expensive than GPT 5.6." This is a flatly false statement.
It may not be false but may be a "category error" [0]. Reserved GPU pricing & bulk inference pricing is 3x to 6x cheaper than "API rates", but renting your own GPU cluster (in this crunch) to run a 600b+ open weights is going to be "more expensive than GPT 5.6".
Even then, it remains to be seen if Huawei will pull their weight (and match up to Nvidia) as spectacularly as their fellow Chinese AI Labs have. If so, the WAICO alliance is ready to go all-in.
[0] Ben, and probably other "influencers" in this space, may be prone (knowingly or unknowingly) to favour points that make their conclusion for them (https://en.wikipedia.org/wiki/Motivated_reasoning).
abernard1 3 days ago [-]
Fair. Too strong a statement.
But much like Ben's point that commoditization is a relatively novel concept to many in tech, it's not the consumer AI applications at risk of commoditization. They have distribution there.
It's the literally millions of engineers who are updating codebases with tools replacing workers partially or wholly. It's the supply-side where there's compression, and no need for distribution.
I would argue, given the enormity of the existing SaaS stack and how it integrates with the human machinery of personnel, that's where volume is. And that is clearly cheaper and a home run.
Commoditizing a ~$100B AI consumer market is no small feat. Commoditizing 20% of the $500B SaaS market, to say nothing of the underlying systems in the who-knows-how-many trillions "Big Tech" market (you're obligated to say that like the Kool Aid man), is shocking.
AnthonyMouse 3 days ago [-]
> Commodity market profitability is determined by marginal cost of production. LLMs have marginal cost; traditional software does not.
This is the story for Nvidia/AMD or cloud providers rather than OpenAI.
> With increasing inference as % of total compute, if labs create efficient models -- which they can, because they can create highly optimized models amortized over very high inference loads -- they can be low cost producers, and be competitive at $/task rates
It seems like there would be problems with this on both ends.
For general purpose models, everybody is trying to make them efficient, so you can't win just by being slightly more efficient. You would have to be so much more efficient that you can charge high margins while still capturing the majority of the market so that the high margins get multiplied by the majority of users and the users you leave on the table aren't funding open competitors. Meanwhile everyone else is also trying to improve efficiency, so one misstep and you're behind.
Example of where this can be a problem: You spend a preposterous amount of money to create an efficient model, then someone else publishes a paper with a new technique that gets a similar but incompatible efficiency improvement out of a model that costs a lot less to create. You have now spent an enormous amount of money in exchange for no competitive advantage.
And from the other end, one of the best ways to get efficiency is through specialization. A general purpose model can generate code or summarize a meeting transcript, but a special purpose model can do it as well or better with far fewer parameters and resources. But then you don't have a situation where one huge AI company has The Most Efficient Model, you instead have dozens of specialized models produced by independent sources that are each the best in a given niche. Any proportion of which could have open weights, or have an arbitrarily small advantage over the ones that are.
Moreover, these problems combine: Both the computing hardware vendors and the AI companies want the margin on doing inference, but the more of it one of them gets, the less the other does. If the AI companies were actually getting huge margins then it would be in the interests of Nvidia, AMD, Apple, Intel et al to fund efficient open weight models in the same way they fund Linux. Commoditize your complement. And those models don't even have to be better, as long as they're good enough that the closed models can't charge a significant premium and the margin shifts back to paying for hardware.
resonious 3 days ago [-]
> for frontier work.
I'll agree that GPT 5.6 may well be the best given the above contstraint, but for run-of-the-mill dev tasks (real ones, not benchmark ones), GLM 5.2 still blows every other model out of the water.
Cost per task as a metric is a bit ridiculous because there are so many types of tasks. GPT-5.6 can do some tasks GLM could only dream of, but GLM can do some tasks 100x cheaper and better than GPT-5.6.
didibus 3 days ago [-]
China is working on the whole supply chain though and they're willing to compete on razor thin margins. Just look at EVs. They build great cars but the competition is so aggressive that investing in any one Chinese EV company isn't exactly an amazing ROI.
I could see AI ending up the same way where the customer captures most of the value rather than the companies. Open weight models are what make that kind of competition possible.
3 days ago [-]
marcus_holmes 3 days ago [-]
> Models are not free. Downloading them is free. Running them is not. This has manufacturing economics, not software economics; the idea that they are "free" is an economic category error as it relates to their actual use
I notice that the article, and this discussion, hasn't mentioned or considered local models.
We can already run a low-spec model on a laptop. Because there is demand for this, it will improve and we will get better laptops and better local models. We will also see models being run on dedicated local hardware and called from the laptop.
If I can download a reasonably capable model to my own hardware and run it without paying anyone for either the model or the inference tokens (effectively making models and intelligence actually free once the hardware is bought) how are the Frontier AI Labs going to make any money at all, let alone enough to support their vast valuations?
nunez 3 days ago [-]
This will matter A LOT more after Apple gets serious about integrating AI into the OS.
marcus_holmes 3 days ago [-]
Agree, all those consumer ChatGPT accounts will be gone
gmerc 3 days ago [-]
Yea but at those rates VCs will never make their money back. Because Deepseek and friends keep releasing the inference optimisations to everyone instead of holding them back to pay their investors.
StopTencent 3 days ago [-]
[flagged]
elictronic 3 days ago [-]
This argument starts falling flat when AI companies literally stole from every business and country on Earth. China’s government can f right off, but so can this failed argument.
tristanj 3 days ago [-]
I did read the article, but it misses the core issue entirely, and it's why I shared my comment to begin with. Look at the cost-per-task benchmarks from Artificial Analysis https://artificialanalysis.ai/models?cost=cost-per-task
Anthropic’s API pricing is getting impossible to justify. Anthropic previously had the highest quality models, and used their position to charge premium prices, enjoying inference margins of over 70% [0]. They could charge these prices because no other model came close.
But over the past month, the market has shifted dramatically. Over every single performance tier, Anthropic is being squeezed on price.
* Low end: DeepSeek V4 Flash runs at ($0.02/task), Xiaomi's MiMo-V2.5-Pro at ($0.03), and Haiku at ($0.24). Anthropic is ~10x more expensive than the Chinese open-weight options.
* Mid tier: Claude Sonnet 5 ($1.53/task) is nearly 50% more expensive than GPT-5.6 Sol ($1.04), nearly 2x the cost of GPT-5.6 Terra ($0.82), and 3x the cost of GLM-5.2 Max ($0.47). There is basically no reason to ever use Sonnet 5, the competitors are significantly cheaper.
* High end: Opus 4.8 ($1.80/task) and Fable 5 ($2.75) are the two most expensive models, and GPT-5.6 Sol ($1.04) and Kimi K3 ($0.95) offer comparable performance for significantly less. Less the fact that Kimi K3 will get ~10x cheaper once its weights are released and served on neoclouds with Nvidia hardware [1].
OpenAI priced their latest GPT-5.6 models cheaply in order to regain market share. When Anthropic clearly had the best models, their 70%+ inference margins were defensible. But today they are the most expensive option in every single tier. Unless they make significant price cuts soon, they run a serious risk of bleeding market share.
[1] "American companies such as Modal, Fireworks, and Baseten will be able to serve Kimi K3, at one-tenth the cost of their Chinese competitors because they have access to advanced Nvidia hardware"https://x.com/rohanpaul_ai/status/2079027313455550839
isodev 3 days ago [-]
Imagine how cool it would be if actual competition prevents Anthropic or OpenAI from becoming an Apple/Google kind of cartel. I don’t care if it comes from China or not.
AlexCoventry 3 days ago [-]
[flagged]
culi 3 days ago [-]
And here I was believing the Red Scare was for the history books
Using a Chinese LLM will not put a Marxist under your bed.
Did you know Gemini is shockingly bad at French poetry? Hasn’t stopped me for using it for all other tasks though.
bornfreddy 3 days ago [-]
> Using a Chinese LLM will not put a Marxist under your bed.
Interesting tangent - it might be taught to introduce stealthy backdoors in your company though. Maybe even across multiple PRs where each session puts a small chink in the armor, and together they allow unlimited access to the attacker who knows about them.
After all, LLMs are mostly black boxes. How comfortable would you be running a Chinese compiler?
bspammer 3 days ago [-]
If that happens one time at any company, it would be discovered and reported on eventually, and people would never trust any Chinese model again. The motivation for China is clearly to prevent America gaining a monopoly on AI, not spycraft. They have much more effective ways to do the latter anyway.
bornfreddy 2 days ago [-]
It can be done so that it looks like a series of "oopsie" bugs. China (the same as US) has used spycraft extensively in the past.
sevenzero 3 days ago [-]
Whats different to US models then? They can also be taught to introduce stealthy backdoors. Its not like stealthy backdoors are a Chinese only topic.
w4yai 3 days ago [-]
NSA has been caught doing that more times than China's MSS
bornfreddy 2 days ago [-]
No difference - very similar actually.
AlexCoventry 2 days ago [-]
Power is more decentralized in the US, even now with Trump.
killingtime74 3 days ago [-]
Good old FUD
solumunus 3 days ago [-]
The valuations are unjustified even at the prices they’re charging now.
They’re going to try their best to offload these investments into our pensions before the inevitable crash.
janalsncm 3 days ago [-]
Right, but retail investors weren’t supposed to find that out until after the IPO.
techpression 3 days ago [-]
Apparently it’s already happening to a degree, wether it continues or not (or even is relevant) is not really my area of expertise.
I had plenty of gains by holding investments which Goldman Suchs were making doom statements about.
rootsudo 3 days ago [-]
Always buy opposite, buy low and sell high.
Too many people here buy high and sell low.
When Goldman makes statement, half the time it is prepped by an associate or two that has minimal experience and goes through a MD that enjoys the gloom and doom. That’s why they publish. Goldman makes money on both sides of a trade.
eru 3 days ago [-]
As an advanced technique, you can also sell high and buy low.
lostmsu 2 days ago [-]
Sounds like there might be some hedging strategy involving both like a box spread!
eru 2 days ago [-]
You are more creative than me. I just had short sales in mind.
rvz 3 days ago [-]
> The people who are most afraid of Chinese models are the VCs who poured into Anthropic and OpenAI at astronomically high valuations.
Correct. These chinese labs has proven that having just the model is not a moat, and the safety concerns were all just attempts at regulatory capture.
This is why labs like OpenAI and Anthropic are panicking and are racing to the exit before their valuations start being questioned.
eeiei 3 days ago [-]
[dead]
lorecore 3 days ago [-]
Good. Over the past few years, VCs have proven that they’re warmongering psychopaths. Hopefully China puts every last one of the Palantir/Flock/Anduril class out of business.
happypappy123 3 days ago [-]
[flagged]
lorecore 3 days ago [-]
I’m 100% certain that China won’t be sending any goons to my front door.
Not that I'm anti-China, but their companies would have no qualms selling surveillance tech to your local gov too if they could. I'm sure Palantir, Flock, and other US companies would sell to China too if they could.
yonaguska 3 days ago [-]
This is true, but there is another foreign country that can send people to your door. What's to stop China from eventually buying that type of influence over our govt officials?
Revanche1367 3 days ago [-]
If our govt officials are willing to betray their citizens for money to China, what’s the point of preferring to give them power over us instead of giving it to China?
platevoltage 3 days ago [-]
Given the nations that have actually bought influence over our government officials, China would be an upgrade. I doubt the president needs another decked out 747 though.
wordpad 3 days ago [-]
I think everyone understands models will be a commodity.
Its the user base (with ads and upselling) and proprietary wrappers which will make money for typical customer.
Even enterprise customers arent going to be spending a lot on tokens. Once labs no longer have to subsidize trainings tokens costs will drop 10x and once models get burned on chips costs will drop 10x more and you physically won't be able to burn significant number of tokens unless you're deliberately trying to.
skeptic_ai 3 days ago [-]
[dead]
testbjjl 3 days ago [-]
We both know the answer. Write offs. If your fund was not in AI heavy you’d have no investors.
fnord77 3 days ago [-]
I guess I shouldn't try to buy shares of OpenAI on the private market...
ProAm 3 days ago [-]
tbh they are floundering even to regular investors. They are trying to give the US gov 5% of the company so they become 'too big to fail' but they are in trouble.
walrus01 3 days ago [-]
I would say it's not just the VCs but the various other entities that will be left holding the bag of debt if the AI-fueled datacenter construction boom/bubble pops. For a list of large and well known projects and their scales:
Aren’t we the ones giving money to the VCs in the end of the rain cycle ?
cyanydeez 3 days ago [-]
mmm, the chinese models are also working on local GPUs at consumer grades. so theyre not just drainig cloud moats.
vrm 3 days ago [-]
good luck running a 2.4T model on any local hardware. it’s not gonna happen. the arrow is to specialized hardware at least for the smartest models
Flere-Imsaho 3 days ago [-]
Yes but someone who has access to that kind of hardware can distill down to a smaller model that is specialised for a specific task. I don't need my local model to be an oracle for everything, I want a coding AI, one that knows medicine, another that recognises objects in my security camera, etc.
cyanydeez 3 days ago [-]
sir, I'm not running a multi billion dollar code base; I just want my nose wiped and a clean fork of whatever repo might be the target of supply chain attacks, and a few nicissities.
I don't need 2.4T to do that; I'm doing it with 35B or 27B. If they get me a model in ~80B with a A5B or A7B, that will be the end point.
It's bizarre people, by themselves, believe all these parameters are getting them much more.
Lets be serious: if we as a civilization really wanted the advancements promised, we'd find the 1000 best scientists and give them free access to these models while the rest of us get personal GPUs for specific use cases.
But instead, we have to endeour this penis measuring contest for the infinite bikeshedding of the universe.
matheusmoreira 3 days ago [-]
I have hope it'll happen one day, even if not now.
nekusar 3 days ago [-]
Already is possible. On a machine with 32GB ram, and NO gpu. Just need a large SSD or NVME. Streams from disk to memory.
But its also a 800B sized model running on a ram constrained system with no GPU.
Techniques, GPUs, more ram, and faster disks can always speed it up. But the point being is they run on low end machines now. Its now an optimization problem, not a possibility assessment.
NekkoDroid 2 days ago [-]
Isn't it more s/tok?
matheusmoreira 2 days ago [-]
That's fucking awesome!
airstrike 3 days ago [-]
they will likely suffer enormous real losses too, not just paper, though not as enormous
for VCs, breaking even is losing
onesociety2022 3 days ago [-]
But there a ton of other VCs who poured money into SaaS businesses. They have the opposite incentive. They want tokens to be cheap like a commodity so the value accrues in the SaaS/app layer.
hamandcheese 3 days ago [-]
Cheap tokens only benefits SaaS that depends on AI. Otherwise, cheap tokens means it is only more cost effective than it already is to cut out the SaaS and build instead of buy.
Ericson2314 3 days ago [-]
Yeah maybe *both" sass and training companies are wiped, imagine that!
pishpash 3 days ago [-]
Maybe AI is deflationary, in the right hands.
eru 3 days ago [-]
Increases in productivity are deflationary, all else being equal.
(And that's good, that's how it's supposed to work. See eg prices for transistors or hard disks or solar power over the last few decades.)
21asdffdsa12 3 days ago [-]
I for once welcome the donations to the public of our generous basilisk worshiping overlords
jillesvangurp 3 days ago [-]
I don't think it's that black and white. OpenAI and Anthropic are building valuable tools on top of their models. You get a very rough and much less polished version of that with open source tools and and open source models. And you still need inference infrastructure to run those. But at this point most of the competition is in the tool ecosystem, not the models. And while there are plenty of people toying with things like opencode there's a clear pecking order emerging where Codex and Claude Code/Cowork are generally considered the top choices before tools like MS Copilot, Gemini, and then a rapidly shrinking long tail of alternatives to those.
In the end what companies pay for is not tokens but results. A DIY kit of models, mac minis or whatever, and a bunch of poorly integrated OSS tools doesn't solve their problem. For the same reason, people use Office 365 rather than running Libreoffice. And for the same reason things like AWS dominate the market rather than people DIYing their infrastructure together themselves. Most of the money is in polished turn key solutions. Which is what Anthropic and OpenAI offer.
The juicy market here is the enterprise market. That's mostly business users, not programmers. They'll be hooking up all their SAAS tools (which they also over pay for), and other stuff. They'll be paying for boring things like data residency, compliance, etc. And they need access to reliable infrastructure to run all this stuff. They'll want this shit to just work and not to be dealing with a lot of poorly integrated stuff.
Most of the billions invested are being sunk into infrastructure, chip design, and access to resources (land, water, energy) needed to run data centers. A handful of companies now own most of that infrastructure and they also happen to have the top models, researchers, the best tools, and warm customer relations. And they sell access via very convenient subscriptions with high enough limits that people don't have to worry about things like token cost. The game here is recurring revenue from customers that like predictable pricing, reliable quality of service, and iron clad compliance and data security & residency, and quality guarantees. These companies don't want to be chasing model quality and have to upgrade their entire company every few weeks. They want continuity and predictability. Mostly they just pay Anthropic, OpenAI, MS, or Google to take care of this for them. There might be some niche EU players that become a bit bigger. But I don't see a large scale switching to Chinese suppliers for a full polished alternative. The Chinese might give away their models. But I don't think they'll be generating a lot of revenue.
And if you want to run your own models, you'll still need infrastructure to run it. These four companies together with the usual cloud giants control most of that and as well of the supply of resources (chips, data centers, energy, etc.) in the EU and US markets. There's going to be a long tail of self hosted and gobbled together stuff but it's going to be a much rougher experience for end users and it won't likely be most of the market any time soon.
phendrenad2 3 days ago [-]
VCs are just pass-through investors, the money comes from billionaires. And when billionaires face losing money, the whole system re-arranges itself to stop that from happening.
sciencejerk 3 days ago [-]
Yep! Maybe Chinese models are banned in USA and made completely illegal
ineedaj0b 3 days ago [-]
Not sure most of money is from VCs.
hnfong 3 days ago [-]
Well, then let's hope you're wrong and the AI bubble won't also blow up private equity and wipe out people's retirement funds...
reinitctxoffset 3 days ago [-]
[dead]
alex1138 3 days ago [-]
I'm a civilian, not a VC. In my own case, I'm worried how many things pass through the CCP. How censorship of mentions of Tiananmen Square is something they're quite interested in
riskd 3 days ago [-]
How exhausting.
alex1138 3 days ago [-]
Why? Why is this not a concern?
JSR_FDED 3 days ago [-]
Because you’re not that important. I don’t mean that in a mean way, just that if you’re a serious player in national security or something like that, you’re already not using Gmail, let alone ChatGPT.
sciencejerk 3 days ago [-]
You would be surprised...
platevoltage 3 days ago [-]
Claude refuses to call Trump a Fascist.
alex1138 3 days ago [-]
You're hilarious. That isn't a compliment
platevoltage 2 days ago [-]
So you only have an issue with Chinese Ai not calling a spade a spade?
nateburke 2 days ago [-]
Releasing open weights that can approach frontier level intelligence (irrespective of number of tokens burned) is just a way of telling the world that anyone, even China, can serve frontier level inference if they have the chips and warm shells to do so.
What is stopping China from gaining a majority market share, then, in terms of serving inference?
AI Sovereignty -- yes
Cybersecurity concerns -- yes
Latency -- no, unlike previous emerging IT workload types , inference does not have strong latency requirements. eg 1s of additional network latency doesn't matter to a 15 min, 10-turn agent session.
Cost -- ultimately this comes down to a nations ability to plug chips into warm shells. which forks into geopolitical / trade on the chips side and energy scalability and modularity on the warm-shell side. Even if you call geopolitical / trade a toss-up, China has the US beat HANDILY on the energy front, yearly they are deploying 10x power to their grid relative to the US, which is shooting itself in the foot at every possible moment.
IMHO chip tech will travel across borders, absent a breakthrough in analog inference, energy scalability will ultimately dominate.
faangguyindia 3 days ago [-]
I operate an analytics site (pretty big one B2B where client's backend feeds data into our system), and we see tons of traffic originating from northwestern China (Xinjiang) from Shenzhen Tencent Computer Systems Company Limited.
There are also half a dozen other companies from China continuously hammering our clients’ websites.
I was wondering, what's in that cold dessert? Low and behold satellite imaging shows massive datacenter build outs, very cheap solar energy.
Few months ago something happened and the Geo location on data on those IP now shows "Shanghai" or "Shenzhen". A way to cover tracks? But mapping latency still points to fact that nodes behind these IPs are still operating around Xinjaing region
credit:
'You Can't Cheat Time: Finding foes and yourself with latency trilateration' https://youtu.be/_iAffzWxexA
HN user: lopoc
Shenzhen vs Xinxiang is hard to do using this technique but Shanghai vs Xinxiang does show difference.
Assuming that China only distills is a huge mistake.
It’s no longer some backward place that does low value copying. Look at companies like ByteDance and Xiaomi.
Chinese companies aren’t just distilling, they’re acquiring data in the same way American companies did by paying people and crawling the internet.
The way I understand it, China has a few large companies that crawl the web at a rapid rate and build corpora. The government essentially wants select few companies to do this and then make the data available to other strategic companies operating within China.
Then there are data aggregators that buy data from apps, websites, and services, as well as systems like OpenRouter or Cursor, where companies can learn from the “traces” of coding agents, chats, and so on.
This massively reduces costs, as smaller companies like DeepSeek don’t have to do their own crawling or acquire data from 100s of websites and coding agents etc....
There are also companies in China that buy American LLM APIs and proxy them to companies within China. So, there could be 10,000+ companies using American AI products, while China logs all of this, understands how they’re being used, and trains on their traces.
feisuzhu 3 days ago [-]
And, non-state run Chinese companies are just like companies in US, they usually don't joint forces to maintain a common infrastructure, if they can build moat (or at least be in leading position for a period of time), they do it, sharing crawl dataset is no go.
feisuzhu 3 days ago [-]
I thought you got the location from BGP registration, then IMHO the before/after are both correct, it might be datacenter in Xinjiang belongs to Tencent.
prox 3 days ago [-]
What does this line mean “ the web at a rapid rate and build corpora.” , what are they trying to do? Suck up data for training AI or something else?
wxw 3 days ago [-]
> It’s striking the extent to which Claude Code and Codex are proving to be quite sticky; whichever harness you start working with is likely to be the one you stick with, and that figures to be even more the case with non-technical users.
My experience has been quite the opposite. I was using Claude Code almost exclusively this winter/spring and swapped to Codex earlier this summer. It took no time whatsoever to switch. And before Claude Code, I was using Cursor. Same story.
[edit: Oh and there was also a brief interlude with Conductor, though I think they're more or less just serving the underlying Claude/Codex harness]
Aurornis 3 days ago [-]
For personal use I agree.
For companies, these decisions are very sticky. Companies go through a lot of red tape to get anything purchased and approved, then they discourage change because it's a lot of work.
So the product that gets a foothold in a company sticks for a long time.
Then a couple years later a sales person convinces an exec that they can save some money by switching, so the switching game begins. Not necessarily motivated by the better product, mostly the price. My wife's company keeps switching their tools out from under everyone every year or two. Just when they get everything stabilized and everyone familiar with the new tool, some new contract is signed that moves them all to some other company's suite.
exhilaration 3 days ago [-]
I'm confused, I work at a big giant Fortune 500, we all get GitHub Copilot subscriptionsn
- we can switch between OpenAI and Anthropic models with just a click in Visual Studio. There's no stickiness at all. They just made us go through a training after the price hikes about how to choose between models for the best cost/benefit ratio.
sothatsit 3 days ago [-]
The models are not what is being discussed here, it is the harnesses. That is, Claude Code, Codex, and what you use, GitHub Copilot. I suspect there would have to be strong reasons for your Fortune 500 company to switch away from Copilot.
Similarly, I have made no ground in arguing to try to get Codex at the company I work for, which got Claude Code a year ago and sees no reason to go through the whole process of setting up any alternatives when Claude Code already works and is at the frontier.
byzantinegene 3 days ago [-]
claude code is free to use, you can use claude code with any non-anthropic model, so harness stickiness doesn't benefit the AI labs at all.
dannyw 2 days ago [-]
Claude Code is closed source software that has had quite a few documented bugs with degradation when using non-Anthropic models, FWIW, I would not suggest using it.
Aurornis 3 days ago [-]
GitHub Copilot is the sticky product in your org.
You can choose a selection of different models within it, but you're not using Codex or Claude Code.
Maxatar 3 days ago [-]
Not sure about Codex but Claude Code works with Github Copilot and can even work with open source models like Kimi/Deep Seek.
2 days ago [-]
tesnorindian 3 days ago [-]
We too use GH Copilot in our enterprise. OpenCode works flawlessly with GH Copilot subscription if you don't prefer VS Code.
thaeli 3 days ago [-]
Same here, except we just defaulted everyone to Auto and expect the percentage of tokens spent via the auto router to be high.
hhh 3 days ago [-]
GitHub Copilot and Visual Studio are the sticky products for your company.
rohansood15 3 days ago [-]
Companies have learned their lessons on stickiness with cloud providers. Every enterprise has a multi-provider strategy now.
andersonpico 3 days ago [-]
Every company that I've worked with that provided models internally did so through LiteLLM and offered both Anthropic and OpenAI models so it was trivial to switch between them.
blfr 3 days ago [-]
Most companies just get you a Claude team sub and maybe a couple of skills.
stingraycharles 3 days ago [-]
We only get Copilot. I’m not very happy.
AgentME 3 days ago [-]
What do you find worse about it? I've been switching between it, Codex, and Claude Code to try to compare them, and my only conclusion so far has been that it's nice that Copilot has both OpenAI and Anthropic models as options.
linkregister 3 days ago [-]
The parent poster is almost certainly talking about inference within workflows and not for interactive coding agents.
Aurornis 3 days ago [-]
I like all the different comments in this thread saying that most companies do X, where X is a different answer from each person: LiteLLM, GitHub Copilot, Claude Code.
HaloZero 3 days ago [-]
I imagine the play here is going be connectors. Can you get slack to avoid integrating with anyone other American ai providers, same with Google suite, etc etc.
linkregister 3 days ago [-]
It's almost trivial to create a custom Slack application wrapping your desired harness running in a container on your organization's k8s cluster. Likewise with MCP support. These are already open.
IAmGraydon 3 days ago [-]
Same here. I flip flop between them. Most people I know who have access to both, technical or not, are doing the same. They’re just too close and sometimes one does what you want better than the other.
nl 3 days ago [-]
Have you ever worked with a non-programmer and helped them setup their AI workflows?
You install MCP connectors, specific skills, work around model/harness quirks, set security boundaries etc.
It's a lot of work, and most people will never want to change it once they have it working.
favouritemartin 3 days ago [-]
Skills are quite interoperable, and you can easily ask Codex / Claude to help you with switching the MCP connectors or any other things specific to your previous workflow. It's been quite low friction in my experience.
nl 3 days ago [-]
I know someone who runs AI training.
They will have people who don't understand the distinction between visiting Claude.ai and downloading Claude Cowork.
They type the words "setup MCP" into Claude.ai and expect it to automate Excel on their machine.
There's a pretty big gap between the things we talk about here, and where the world is at.
andrewf 3 days ago [-]
It strikes me as like setting up an IDE. People have preferences, switching is possible, but there are advantages to saying "we are a Visual Studio + Resharper shop" or "everyone uses IntelliJ to work on this project".
trollbridge 3 days ago [-]
Yes. I taught the non-programmer to ask the harness to set up things like MCP connectors.
cyanydeez 3 days ago [-]
we have AI. WHAT is it good for if a harness cant just take a api endpoint and some permissions and duplicate.
its so distracting seeing these types of confision.
every plugin is already just multimodaling their targets.
solumunus 3 days ago [-]
I think they stickiness is less about the difficulty of switching and more about the lack of desire. I’ve been using Claude since day one, it works well and I’m happy, I like it. I’m sure Codex is good too. Switching from one to the other certainly isn’t going to be a game changer, the discourse shows me the differences are marginal.
Probably the only reasons I would seek change are economical.
jjfoooo4 3 days ago [-]
A sticky product is one that switching away from creates a major hassle. Which means the user will pay more to avoid said hassle.
“I don’t really have a strong preference between the two” is another way of saying “the product isn’t sticky”, which is another way of saying “this provider has very little room to increase margins”
solumunus 3 days ago [-]
No, I don’t think so, and searching seems to confirm my view. Inconvenience of switching is just one aspect.
There’s little difference between Coke and Pepsi and the barrier to switching is nil, yet clearly the products have stickiness. People have slight preferences and become familiar with the brand and then engagement becomes habitual.
The effects on margins are irrelevant to this.
mediaman 3 days ago [-]
Convergence in coding makes them highly substitutable. But I could see harnesses configured for different purposes -- let's say, a harness for creating teaching plans -- being able to cater to its audience better than a coding harness. Maybe it's got tools to plug into standardized curricula, what the lesson books will be, what other lesson plans the district's teachers have made, etc., which could be done in a clunky way in a regular harness but could be streamlined.
bushbaba 3 days ago [-]
agreed, my F500 company switched off claude code to copilot in 30 days. All 5k+ engineers. That is the fastest migration i've ever witnessed. This includes switching all our agents from Claude SDK to Copilot SDK.
sergiotapia 3 days ago [-]
My same progression here. I started with ChatGPT website, then Anthropic website, then Cursor, then Windsurf!, then claude, then opencode, then ohmypi, then codex, finally back on Cursor now because I think they cracked the UX for what great dev looks like. The grok 4.5 fast model + cursor ergonomics is insanely good!
The cost of me moving around these different AI models and harnesses was pretty much 0.
Aperocky 3 days ago [-]
I use a mixture of claude code and codex and kiro as my swarm.
They communicate through my own harness, and it's working pretty well so far. claude code is being overtaken by codex however because I noticed lately the accuracy of the latter is the best.
philstephenson 3 days ago [-]
As a Hacker News user and commenter, you are not the type of user he’s referring to.
happypappy123 3 days ago [-]
[dead]
SOLAR_FIELDS 3 days ago [-]
Which would imply that these things are fast becoming… checks notes… a commodity?
WinstonSmith84 3 days ago [-]
I'm rather scared of US models - if Anthropic was the only AI provider in the world, it's easy to see that common people would have no access at all. Thankfully there is OpenAI which compete 1:1 with Anthropic (at a slightly lower cost) but most importantly the Chinese models keep Anthropic, but also OpenAI in checks.
And I'm saying this as someone working for American companies.
kiicia 3 days ago [-]
it's a real paradox, that chinese models are what guards democratization and private use of ai while us models are moted castles with "kings" crying that you are stealing their legally stolen goods... what times we are living in...
maerF0x0 2 days ago [-]
It blows my mind that these two companies are going to mint billionaires, meanwhile https://en.wikipedia.org/wiki/Aaron_Swartz was bullied to death by the state for sharing academic papers which were at least in part, if not largely, paid for by tax dollars.
marcosdumay 2 days ago [-]
> which were at least in part, if not largely, paid for by tax dollars
They were all in the public domain too.
davidpapermill 2 days ago [-]
I'm afraid law and morality are two separate concepts. You make a good point.
dofm 2 days ago [-]
"You need to understand that Sam can never be trusted ... He is a sociopath. He would do anything."
kingleopold 3 days ago [-]
ethics and other details are for humans. AI companies just proving it, even highest IQ teams are against ethics because they want more, even several millions is not enough for them.
jml78 2 days ago [-]
AI has forced a dramatic shift in my view on intellectual property. The laws as they exist only protect corporations now.
They do not protect individuals no matter how much people want to think they do. AI has proven this.
I think to level the playing field all copyright, trademarks, and patents laws should be eliminated.
If I want to make a marvel movie, I should be allowed to and profit from it.
AI let the cat out of the bag and there is no going back. We need to let individuals profit just like corporations can from what is considered theft right now.
_DeadFred_ 2 days ago [-]
Or we could build on and enforce the 300+ years of thought and trial that went into copyright law so that individuals are re-empowered.
But nah, let's do the most radical, least thought out thing, and absolutely destroy small scale creators. I'm sure Amazon will be benevolent and continue to pay writers in your scenario.
What is with 2026 and just conceding civilization to the worst actors, and then adopting the worst tactics/thoughts/concepts?
projektfu 2 days ago [-]
Writers will be forced to go on tour like musicians in order to get anything from their work, and they may not release the works to the public until after their death. So much for the last of the troubadours.
marcosdumay 2 days ago [-]
When exactly did rulers thought about copyrights with the goal of empowering the people?
fidotron 2 days ago [-]
> AI has forced a dramatic shift in my view on intellectual property. The laws as they exist only protect corporations now.
Even this isn't quite right. Publishers unambiguously have been screwed.
What has happened is if you can give the political/investment classes enough upside opportunity in the entities doing the IP theft then you're golden, and they increasingly have no problem with even pretending to hide it.
economistbob 2 days ago [-]
You are on to something. The extensive copywrite terms have ensured that the work of everyone except the oligarch corporations vanishes into the ether since it cannot be reproduced for over a century if the author lives another 30 years.
There will be no one debating the great minds of the twentieth century because the corporations made it illegal for people to republish critical editions of any such works.
The greats were recirculated every twenty or thirty years for centuries.
The smaller authors will vanish into the nothing and Western Civilization's 20th century onward will vanish and be almost forgotten forever because of that mouse. There will be more ancient literature remembered than literature after the invention of the printing press because they made it illegal to share printed materials for nigh on two centuries if an author was young when they wrote it. There will be more ancient scrolls preserved for the future than books of 20th century philosophy and science because the latter is a crime.
Only the oligarchs who pirated the books will have a cultural memory. They have cursed our era to oblivion when it comes to intellectual property because they can generate billions now for the mouse's henchmen.
inigyou 2 days ago [-]
The technique is called "accusation in a mirror". By accusing the enemy of doing what you are doing, when they call out what you are doing, they look like they are just weakly repeating your own accusations because they don't have any truth. And the anger that should be directed against you (because of your practices) gets directed at the enemy.
desterothx 2 days ago [-]
Well implemented by the current US administration. Every accusation is a confession
BaseBaal 2 days ago [-]
The US distilled this from the Russian model.
hnfong 2 days ago [-]
There's no way anyone could have made a model so good, it must have been distilled from the future.
Lutger 2 days ago [-]
Thank you for this, I've noticed this often in politics and didn't know it has a name. Some politicians do this so often, that any accusation to me reads a confession instead.
hnfong 2 days ago [-]
It's usually called self projecting in more common terms I think.
TremendousJudge 2 days ago [-]
why is it a paradox? guarding their IP overseas has been the modus operandi of American software companies since their inception
mihaic 2 days ago [-]
The paradox is that their IP was obtained by distillation of the collective IP of humanity.
WarmWash 2 days ago [-]
I'm sure all those stackexchange experts will be getting their checks for helping launch countless software projects over the last 20 years any day now...
mihaic 2 days ago [-]
You're basically saying that since open source projects can be reused by others, the people making closed source projects shouldn't complain when the same thing happens to them.
Everyone posting on Stack Overflow knew it would be public and free. Book authors put an explicit copyright notice, which includes derivations of their work.
kiicia 2 days ago [-]
It’s not their IP, they never created data they were training on in the first place, at best they derived existing data without asking if they can commercialized it
spyckie2 2 days ago [-]
Competition is a good thing for consumers.
Chrisszz 2 days ago [-]
It is also good for the technology itself, just image if there would be just one company that does "just well enough", they will be unmotivated to improve their ai models
mapt 2 days ago [-]
The Chinese models keep American financial markets at "To the Moon" levels rather than "Igniting Jupiter as a binary star" levels. This amount of capital pressure exerts its own gravitational pull in markets and in geopolitics, and every additional dollar of valuation can propel acquisitions, which propels valuation, et cetera. Undiluted, very quickly Amazon owns countries like it today owns county governments.
persedes 2 days ago [-]
this has been my tinfoil hat theory. Investors in US models might be supporting "open" models as a means to create FOMO for other investors to supply more cash to US model providers, which in turn increase their investment value. Package it so they can beat those "adversaries", equate that success with global power struggles etc etc. Seems to be quite effective.
chrisss395 2 days ago [-]
Re: common people, yes, and given all the talk of job instability AI is creating, why hasn’t Anthropic or OpenAI offered discounted plans for laid off or displaced workers? Wouldn’t this be a tangible way to display goodwill and build adoption?
stronglikedan 2 days ago [-]
> why hasn’t Anthropic or OpenAI offered discounted plans for laid off or displaced workers?
Because it's not a widespread phenomenon. A few large tech cos laid off large swaths of people a handful of times. That's only happening in those large tech cos. Most cos are empowering their employees with AI as a tool, not a replacement, and they're not letting anyone go (unless they refuse to use this new tool).
Of course, they're not hiring as much either, since their current teams can accomplish more with AI as a tool. Maybe the AI companies could give job seekers a bit of a discount, but that would be abused to all hell without crazy administrative overhead costs, so why would they?
gowthamgts12 2 days ago [-]
> why hasn’t Anthropic or OpenAI offered discounted plans for laid off or displaced workers?
because they are profit maxing. the AI world would be a completely different place if it's not for the open models.
monooso 2 days ago [-]
I'm sure this will be a terribly unpopular opinion, but I've long held the view that a Chinese AI company might screw me over at some point for a complex geopolitical reason that I don't fully understand. An American AI company will screw me over tomorrow for a quick buck.
As to the argument is that (only?) the Chinese labs are training on my data, I find this almost comical given the amount of highly-personal data companies such as Meta and Google have been harvesting for decades.
over_bridge 2 days ago [-]
I don't trust any American company because they will switch into extraction mode eventually and squeeze every cent they can out of you.
Even if the founders didn't want that, eventually the upper ranks will fill with MBAs and the board with private equity and they will make it that way. Their bonuses are based on quarterly performance not customer experience
Long gone are the companies who served their communities for decades or centuries, providing a stable return to the owners, jobs for the workers and value to the customers.
I'd rather China have my data than America. China might do something with it one day but America built exploiting it into the business model (disclaimer that I'm not Uyghur or Taiwanese though)
subsistence234 2 days ago [-]
the article isn't saying you should be scared of chinese models
braebo 2 days ago [-]
Saying GPT 5.6 Sol is 1:1 with Fable is laughable. OpenAI models are braindead in comparison and I have hundreds of hours with both.
deepvibrations 2 days ago [-]
No doubt Fable can outperform on many coding tasks, but the way you phrase it really exaggerates the gap and I'd suggest that for 95% of tasks, most devs simply don't need Fable level, and in fact, despite mostly using Claude, I have found codex is generally quicker at most tasks and can be even better than Fable with certain languages.
aroman 2 days ago [-]
You have hundreds of hours with a model that was barely even released hundreds of hours ago?
The perception of capability varies greatly between task. For my needs for example sol xhigh consistently outperforms fable xhigh.
jayd16 2 days ago [-]
You could run hundreds of agents in parallel and arguably that counts.
aroman 2 days ago [-]
No, it wouldn’t. The hours in question are human experience, not that of the agent.
jayd16 2 days ago [-]
Why not? You have far more results to review.
If you run a model on slower hardware are you getting more experience? Surely its a factor of model output reviewed and not human time.
aroman 2 days ago [-]
Right, it's about how much time the human spent - the time spent by the machine itself is irrelevant. As you rightly point out: that is why we measure programmer experience by wall clock time, not CPU time :)
aaronrobinson 3 days ago [-]
“ Let the frontier labs win by being better; don’t let them define safety or security, or pull up the ladder of humanity’s collective knowledge”
Love this.
hahahaa 2 days ago [-]
Try different harnesses people! I am actually preferring Chinese models at a fraction of the frontier price for coding. Yeah you need more tokens per unit of work done, but it is way cheaper still. Using CC/Opus as a staff eng / frac CTO. And Hermes/Chinese model as hopefully my team of mid levels. This way I can make good use of pro plan and then get cheap Chinese tokens for the rest and not hit a RL and know it can scale up. plus choosing your model is so cool and some are a lot less verbose.
Hermes is a better coding tool IMO. I can't put my finger on why but it just feels better. Maybe being true yolo helps.
giancarlostoro 2 days ago [-]
That's nice and all, but I would not get hired in many places that are heavily regulated and risk adverse, and would not hire someone who swears by said models because there is no trust in their creators not training for malicious intent, a random tool call here and there, and you've got a "open weight model" that can send your code anywhere.
There's just no trust in a country that is digitally totalitarian and hostile towards its own people. Do people ever look at the full sized Tianamen Square photos? This is not even the photo of the many people on the ground who were killed by their government and it is still insane to look at.
> There's just no trust in a country that is digitally totalitarian and hostile towards its own people.
Are you referring to USA, China or EU here?
whynotminot 2 days ago [-]
What you personally believe and perception within the United States are not the same thing.
preg_match 2 days ago [-]
The US is most definitely hostile towards its own people, and we also have the most comprehensive surveillance apparatus in the world. On par with, if not more extensive than, China.
People largely can't protest here right now, and US citizens are being killed. People are being sent to work camps in countries where laws do not apply. And our leadership is, right now, priming the American people for when they reject the results of future elections. Not to mention everybody involved in the last attempted coup was pardoned - after we were told that our current leadership had nothing to do with the coup.
The state of the US is much more dire than most people are letting on. And it's understandable why. Nobody likes bad news, and we all like to believe things will be okay. All I know is I can't take my phone into the airport. I can't go out and protest without risking my life and freedom. I can't drive anywhere without my location being tracked and logged. And, if I get pulled over, I must comply with any order, no matter how unlawful, otherwise I risk being executed in the street.
Some of these things have been going on for a while, and some are new. But all are real.
buellerbueller 2 days ago [-]
>The state of the US is much more dire than most people are letting on.
This is unfalsifiable, so not a great claim. You can continue to claim this forever without having to prove it.
>All I know is I can't take my phone into the airport. I can't go out and protest without risking my life and freedom. I can't drive anywhere without my location being tracked and logged. And, if I get pulled over, I must comply with any order, no matter how unlawful, otherwise I risk being executed in the street.
Wild, wild exaggerations, but you will point to your unfalsifiable claim to justify them.
The VAST MAJORITY of people in the US who engage in the things that you claim that they cannot, do so, and with no consequence.
I am no fan of this administration, but I am even less a fan of outright exaggeration (which is even one of the defining hallmarks of Mr. Trump)
preg_match 2 days ago [-]
None of this is an exaggeration.
It was just recently ruled that your phone can be searched extensively at an airport, without a warrant. The only reasonable thing to do is simply not bring your phone.
Notice what I am claiming and what I have claimed. I am not claiming these things will happen. I am saying there is a risk. People have gotten executed by the state for not complying with obviously unlawful orders. It has happened many, many times. And before you argue: yes being shot by the police is "public execution by the state without a trial or conviction". That sounds harsh, but that doesn't make it inaccurate.
Every time you are pulled over, there is a risk. Any time you enter a public space, there is a risk that ICE, which is a domestic peace-keeping military force, will arrest you or kill you. There is a risk. There is a risk you will not receive due process. Every time you go to a protest, there is a risk your phone will surveilled, sometimes remotely. We know ICE uses 5G devices to surveil the the air waves. Every time you go to a protest, there is a risk you will be subjected to tear gas or rubber bullets.
All of these things are real, are proven, and have occurred on many occasions that we know of. There is really no dispute here, so you can try to dispute it, but if you do, then you are outing yourself as a dishonest person, and of course then naturally nobody will waste their time talking to you.
Now, you probably aren't happy about this. Neither am I. But emotions such as unhappiness do not magically override reality. Whether these things are happening or not and whether they are a risk is a separate question. A question with only one reasonable answer: yes.
buellerbueller 2 days ago [-]
>The only reasonable thing to do is simply not bring your phone.
Actually, not bringing your phone is eminently unreasonable, particularly to an airport. Sure, you could wait in a line for a paper boarding pass, but if everyone were "reasonable" and did this, what do you think it would do to airport lines? What would happen when a flight gets canceled, and everyone rushes to the desks for help?
Buddy, birth has a 100% fatality rate. Life is risk.
preg_match 1 days ago [-]
If anyone knows that life has a 100% fatality rate, it would be me, with cancer. Trust me I know that.
It’s not an argument, it’s just not.
There is an amount of acceptable risk, you’re right. How do we calculate it? By making sure we follow due process, the law, and we hold parties accountable.
ICE is allowed to execute Americans. Yes, allowed. Because they’ve done it many times, and face 0 accountability. So they are allowed to do that.
Now use your knowledge about humans. When humans are allowed to do something, what happens? If we let people get away with murder, what is the end result? This isn’t rocket science buddy. Put on your thinking cap.
Ditto for the police. When the police make mistakes and violate your rights, they face zero consequences. Your average McDonald’s cashier faces more consequences if they forget your god damn ketchup. So the police are allowed to violate your rights, yes they are. So they do, obviously.
I’ll admit: this conversation is a little bit frustrating to me. Because this is not a deep analysis or conspiracy. It’s just very simple facts and then some very basic incentive analysis. It’s disheartening that there are Americans going around not putting any thought into the power structures of entities they interact with. It’s dangerous, too. People like you are the best target ICE and the police could ask for. They want people who lollygag and don’t care about their phone getting tracked, aren’t concerned about tear gas and rubber bullets, and who believe their rights won’t be violated.
buellerbueller 1 days ago [-]
I am sorry that you live in so much fear and anger. You have already been beaten.
colinsane 2 days ago [-]
hopefully all of the above?
at some point the conversation has to go past this reflexive "USA uses tech abusively -> but look at how abusive China is with tech -> ..." back-and-forth to acknowledging that neither party is your savior -- and then (hopefully) acting upon and coordinating around that understanding.
hahahaa 2 days ago [-]
And how is US doing on digitally totalitarian and hostile towards its own people?
In practical terms you could get US and Chinese models to review each other, right. Depends what your use case is. Coding is kinda not so bad it is reviewable and immutable/traceable per commit. An AI app that is like a psychologist or something may be more worrying.
AlexandrB 2 days ago [-]
> And how is US doing on digitally totalitarian and hostile towards its own people?
Seemingly better than everywhere else in the world, including where I live (Canada). Whataboutism doesn't really work when you use the least bad option as an example.
FeloniousHam 2 days ago [-]
And whatever can be said about the US, at least with democracy you get a chance to fire the leadership every four years (or even sooner the closer you to the place you live!).
mynameisbilly 2 days ago [-]
Are you familiar with the flock camera situation in the States? The American surveillance state has rapidly been getting worse with the help of this technology.
hahahaa 2 days ago [-]
Whataboutism moniker doesn't apply when the context is "avoid Chinese models because China commits crimes against humanity" and context is also "Top sota models for useful work are created by, let's list them all: US and China". If the argument is don't use Chinese models because China is hostile then let's say "The USA.... Objection: Relevance? Overruled".
US is less hostile in the day to day sense, but they could look at your cloud data at any time because say someone in your company protested and exercised their 1A right, and the government didn't like that, so that sort of thing tracks directly to secrecy and privacy guarantees I might want from models.
I personally don't worry for my sloperating but if I ran a large company's AI I might consider it. Esp if in Europe.
I think in general rest of the world needs to take notice (not saying afraid), starting with the US. It cannot be taken for granted that China's frontier labs will be a few months behind. They might be at par or exceed.
The lessons from steel, solar and EV needs to be learned by all lawmakers. You have to respect and learn from how China Government puts the system in place for complete industry takeover and they have been very good at it. The problem with AI is that democracies will be inherently slow in adopting AI, unless something changes in the system.
At minimum, every democratic Government (US, Europe, India) need to build long-term AI vision and execute that no matter which party comes to power. Additionally, be ruthless about protecting domestic labs. It can only be possible if the intelligence pricing by domestic labs per productive task is in the similar range as open-weights models. Right now, it is not the case, even if the article gives the example of Sol vs K3.
Protecting domestic labs means not bailout, but fast track to cheapest energy, fast track approval for data centers, enforce some guardrails so customers get to use the open weights models only hosted in the country by US (or Europe) businesses. Without these protections, it might be a slow death.
awakeasleep 3 days ago [-]
In the earlier days of the USA we did the same thing, with our government having an industrial policy that fed US industry and put us ahead of Great Britain.
It doesn't have anything to do with the form of government, it has to do with the aims of the government.
Barrin92 3 days ago [-]
What is this panic and protectionism supposed to be good for? This is open source software, there is no "AI industry", there's virtually nobody employed in this. "Domestic AI" makes about as much sense as a domestic Linux kernel. If the Chinese want to subsidize the world's water and energy use to supply the world with chatbots good luck to them. There's no need for guardrails or fast track data centers, they can plaster their entire country with data centers to churn out slop, I'm glad we don't
bg24 3 days ago [-]
I think it is deeper than that. "LLM => peak of productivity" takes way less time and effort than "Linux kernel => any productive work". Compare Dec 2025 vs July 2026 models in terms of capabilities.
Nobody can predict 5 year out. However, the country that can be ultra efficient by making their governance, health, manufacturing, military, etc AI-native will be far ahead in the game.
Barrin92 3 days ago [-]
There's no evidence that these things make anyone more productive, in fact the opposite, empirical research suggests users are less efficient while thinking they're more efficient. You're invoking "AI" the same way the blockchain people did before they all finally died out. "You need to put the government on the blockchain bro, you're gonna lose out to the future bro", turns out you don't
But that's not really even relevant to the debate. Insofar as software has made the world more efficient, it doesn't matter where it's written. That's the point of open source software, there's no tacit knowledge. When a country loses its nuclear engineering capacity that's dangerous because it takes a long time to rebuild. The only reason the Chinese are already competitive on LLMs but haven't managed to make a state-of-the-art jet engine is because the latter, unlike software, is difficult to copy and you don't need to worry about something you can copy.
hnfong 3 days ago [-]
Only those who intended to use the new technology as a tool to achieve dominance and control are freaking out.
And it's the only viable tool the US has left. It's reasonable they are freaking out.
eeiei 3 days ago [-]
Ok…
0x38B 3 days ago [-]
Excellent article; the argument towards the end for allowing distillation for US companies is compelling:
> To that end, here’s an even more interesting question around distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here?
> In fact, this paradox is the solution. I believe that open weight models are good for innovation (and, per the above, I think that labs on the frontier will be fine), but it’s a problem to be dependent on China. The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation, for U.S. companies at a minimum. Stopping distillation — which is literally just querying the API — is nearly impossible; the U.S. should go the other way and lean into a new copyright policy that both indemnifies the labs and also guarantees that what they learned fuels further innovation for everyone else.
protocolture 3 days ago [-]
Very good point.
That would prevent the facebook strategy of sucking up MySpace users and then defending TOS that prevent other social media apps from doing the same to them.
fellowniusmonk 3 days ago [-]
The U.S. "executive" class is so obsessed with the "exploit" part of the explore/exploit cycle that it's very clear they are prematurely closing advancement. Better a little money and power for them now than a lot of money and power for their country/humanity.
This has an element of stochastic improvement so it's hard to predict but the chance of the U.S. "winning" this "race" is pretty bleak.
You see this all the time in communities that have internalized hierarchy as a "good", little kings of shit mountain vying for less and less at a higher and higher cost.
XorNot 3 days ago [-]
My personal hypothesis here is the Chinese government looked at the game and simply decided not to play:
An astute Chinese analyst could reasonably forecast that they had little chance of controlling the AI market due to sovereign trust issues, but would also note that AIs are just software.
When the dust settles the US still won't have factories, and the real value of AI models is still going to be embodying them and getting them to do real, consumer facing work.
Perhaps the most striking thing about the AI boom is how quickly the US abandoned the veneer of local manufacturing in favor of more expensive buildings producing nothing you couldn't make anywhere else on the planet...from imported parts.
_carbyau_ 3 days ago [-]
Yeah, how much of this is China waving distracting AI hands over here while the US Genius-In-Charge watches and completely ignores reality.
throwa356262 3 days ago [-]
According to openAI's own @deanwball:
Even OpenAI isn't buying this distillation talk:
I’m struggling to understand this perspective. Is he using the words accelerationist/decelerationist in a sense other than the obvious one?
EDIT: I searched his twitter history and discovered that his argument is basically “if you drive down costs, then OpenAI will have less money to invest in development, slowing down the overall rate of AI progress.” IMO this take betrays an overwhelmingly stupid degree of exceptionalism, but I guess that’s what I’d expect from someone working at OpenAI.
aesthesia 3 days ago [-]
Nah, the point is that if models are commoditized and there's no hope of making significant profits, no one is going to be willing to make the massive investments necessary to continue pushing the scale frontier. How large a training run do you expect investors to fund out of the goodness of their hearts?
wayeq 2 days ago [-]
> no one is going to be willing to make the massive investments necessary to continue pushing the scale frontier.
that sounds great to me. anything that slows down the pace and gives us a chance to prepare for, at best, massive job displacement, and at worst, robots turning us all into paperclips.
protocolture 3 days ago [-]
The real problem is that I could probably solve even biggerer issues if the investment went to me instead of OpenAI, so OpenAI has a moral imperative to shut down and send me all their money right now.
twelvedogs 3 days ago [-]
what good does it do to invest in a solved problem?
if open models are good enough then it doesn't matter, if they aren't then there's probably a return available in investing there.
lenkite 3 days ago [-]
Doesn't this merely mean that states will do the job of investment ?
eeiei 3 days ago [-]
What a delusional f-wit lol he wants protection of profits for reinvestment?
Every company wants that!
nothercastle 3 days ago [-]
This guy is predicting AI covid escaping from a Chinese lab. I find that kind of silly
So he thinks open weight models will lead to “AI communism” and “dystopian hell” and in the very next point proposes that the US create a federal agency to discourage the use of open weight Chinese models. The motivated reasoning in this post is unreal.
thraway3837 3 days ago [-]
Can you or someone please explain several of the claims made in this tweet?
"I am personally surprised the Chinese state continues to allow the open sourcing of models this good, given potential risks" what risks?
I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). Confused what this means
Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models. Confused again.
One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a "public good" which will ultimately be provided by the state as a kind of "digital public infrastructure." This future strikes me as a dystopian hellscape, but I've never met an open-weight models advocate who doesn't ultimately concede this is where things end. I don't understand this at all.
Can someone in the know please use plain layman's terms to explain what this tweet is about?
slopinthebag 3 days ago [-]
> I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). Confused what this means
I think it's referring to the belief that LLMs are not the path towards AGI, and that LLM's, while useful, are not going to have the impact that the American labs believe it will have.
Barrin92 3 days ago [-]
>Can someone in the know please use plain layman's terms to explain what this tweet is about?
The Silicon Valley people like this openai guy, high on their own supply, are convinced they are building some machine god that will either bring about the end of the human race or utopia, they therefore cannot understand why the Chinese (or any other normal person on earth) are not afraid of chatbots and have other things on their minds.
i2km 3 days ago [-]
It's so bizarre. I keep on wondering where the root of this psychosis lies. Is there some common sci-fi literature that sparks these fears among the readers? Or is it all just latent religiosity and spirituality finding an outlet?
I mean, what an appalling way to live. To take themselves so seriously and simultaneously be terrified of what they're building. If we are really seeing the collapse of the bubble now, I wonder what's going to happen to these clowns when their AI god fails... what will they move on to next?
eeiei 3 days ago [-]
[dead]
whywhywhywhy 2 days ago [-]
> I am personally surprised the Chinese state continues to allow the open sourcing of models this good
Writer seems to have no clue how IP actually functions in China
paxys 3 days ago [-]
> "I am personally surprised the Chinese state continues to allow the open sourcing of models this good, given potential risks" what risks?
I assume they mean the risk of opening up "forbidden" knowledge to the masses without adequate control, which the CCP hasn't historically been known to do.
> I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). Confused what this means
Yann Lecun is a pioneer in the field of AI and Meta's former AI head. He is famously anti-LLM, and considers the entire technology a dead end to achieving human-level AI. The author is saying the CCP has similar views (that LLMs aren't going to get exponentially better/lead to AGI) which is leading them to not control these models as tightly as they otherwise would.
> Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models. Confused again.
"AI accelerationists" = people who want AI to progress. According to the author these people should not celebrate open models because open source = less commerical value in LLMs = less investment into the field (because how are companies going to get returns?), and this will ultimately lead to slower growth.
The last bit is about government controlling AI vs commercial companies. According to the author the former is a dystopian hellscape.
IMO even if you think his points make sense, his job title ("head of strategic futures @openai") means they should all be taken with a massive grain of salt.
mindsolutions 1 days ago [-]
[flagged]
nothercastle 3 days ago [-]
Chinese ai is bad. It’s slowing down progress and it’s so bad we called out the c word and asked for more regulation. Basically advocating for more government assistance to openai
deaux 3 days ago [-]
> This is a point that bears repeating: because U.S. open weight model makers must follow the frontier labs’ terms of service, they (1) are worse than Chinese alternatives and (2) end up distilling the distillation, just with a detour through Chinese labs. Wouldn’t it be better if western open weight model makers could go to the source?
This is of course a baseless assumption. Let's say China created GPT 3.5. Then I can guarantee you that Ben would say "Western frontier labs are at a disadvantage when gathering data, because they have to follow the terms of service of Western media, and Western copyright law". Which we now know wasn't true.
And sure, some will say "but Anthropic can more easily block this as it's a single point of failure". But it's doable to overcome this. Without being "state backed".
jmclnx 3 days ago [-]
One thing I have not seen mentioned between Chinese AI vs US, population.
China has a billion+ people that their AI can "study". Plus due to China's political structure, their AI has access to everyone's chats, comments and sites, scraping everyting.
Here in the US, with 1/3 the population, the AI race was lost before it even began. Plus in the US, all companies and people are doing all they can to restrict AI from scraping sites and peoples chats.
So I believe, China will end up owing AI.
gerdesj 3 days ago [-]
"China will end up owing (sic) AI"
I think you hit the nail on the head - right there!
3 days ago [-]
credit_guy 3 days ago [-]
People who claim that the Chinese open weight models have some type of manifest advantage don't realize that the close weight models have a huge advantage as well: the researchers from OpenAI, Anthropic, Google, xAI, Meta are not dumb, they can read the white papers written by DeepSeek, Moonshot, etc, and they can inspect all those architectures and they can pick and choose the best tricks there are out there, and of course, they have access to their own in-house secret sauces.
Sure, any model that is not at the frontier can use the frontier model to generate synthetic high quality training data, so this can reduce significantly the training costs.
But at the scale of OpenAI, Anthropic and Google, it is quite likely that the (raw) training cost is very high anymore. Here's a few heuristics:
1. All the hyperscalers see a huge demand for inference. They can't deploy datacenters quickly enough to satiate all the demand they see. But, it's is impossible for the inference demand to be constant throughout a day or a week. If you use the times when the demand is lower than the peak demand (which is almost all the time) to dedicate the spare compute capacity to training, then your the cost of training compute is zero.
2. It is likely that increasingly a higher cost of the "training" is actually setting the guardrails, which is essentially post-training. As we've seen, without proper guardrails, the US Government won't allow you to serve inference. Anthropic was hit directly, but OpenAI delayed their 5.6 release as well to make sure the US Government is ok. This part of the training cost can't be reduced easily by using synthetic data generated by other models.
3. The frontier labs are also investing more and more in building an ecosystem around their models.
I am not a frontier lab insider, but take a look at the jobs posted on the Anthropic career page [1]. There are 74 jobs in "AI Research and Engineering" and by my count at most 15-20 are related to pure model training (of pre-training or RL type), and the rest are post-training, safety and security, alignment, interpretability, productivity and lots and lots of other things.
People who claim that Postgres has some type of manifest advantage don't realize that Oracle has a huge advantage as well…etc
killingtime74 3 days ago [-]
If the Google and meta engineers are not dumb how come they consistently trail behind the frontier labs and even the Chinese labs with a fraction of the funding.
Probably bad leadership
hnfong 3 days ago [-]
I always suspect they have the most to lose if legal decisions on copyright issues don't go their way.
Imagine a scenario (theoretically possible but increasingly unlikely) where a US court decides that using "pirated" copyright data to train models is illegal. Now the AI developer has invested hundreds of billions of capital into a thing that is declared illegal and has to be scrapped.
This risk affects existing megacorps more than "startups" like OpenAI and Anthropic (and Chinese companies), because the megacorps have much more to lose. They actually have the cash to pay damages if the flood of copyright claims arrive at the door. This will not only bomb their AI development, but also the rest of their established businesses as well.
And thus I strongly suspect legal issues are holding them back a bit. Megacorps want to win the AI race, but not to the extent they stake the rest of their established business, while the newer companies' only product is AI, so they have to go all in.
Notice for example how Meta's Llama performed much more poorly after they got smacked by a bunch of lawsuits claiming that they torrented a bunch of copyright data.
(Disclaimer: I'm an outsider and everything I base my speculations on is public knowledge.)
kubb 3 days ago [-]
That plus they don’t distill so they have worse RL examples.
shunia_huang 3 days ago [-]
But they both spent tons of money on data collecting/labeling/generation, how is it bad compared to distillation? I thought their data are much better if they spent that much, and it seems they are stupid because with that much of resources putting in there with merely no output compared to the frontier models.
kubb 3 days ago [-]
Creating a RL example by hand is hundreds of times more expensive than generating one using an LLM.
Of course the Chinese companies have incredibly talented researchers, and smaller, better organized org structures which account for the rest of the difference.
BoredomIsFun 3 days ago [-]
they, rightfully so, have no faith in LLMs.
protocolture 3 days ago [-]
>they have access to their own in-house secret sauces.
I remember some feature lauded by Gemini was reverse engineered by the open weights guys in < 30 days.
If they dont publish some technical information its hard to protect in the US, but conversely, once it is published smart people from outside the copyrightosphere can start working to reverse engineer it.
>3. The frontier labs are also investing more and more in building an ecosystem around their models.
Theres nothing there that isnt immediately replaceable.
credit_guy 2 days ago [-]
> There's nothing there that isn't immediately replaceable.
Indeed. But that was not my point. My point is that we still have this old impression that training cost is dominated by compute and it is hugely expensive, and the Chinese labs can short circuit that by distilling the American frontier models. I don't think the training compute cost is a big factor anymore for the American frontier models, because of the reasons I gave. If the Chinese models can get the training compute cost down by a factor of 100, that's not going to make them 100 times cheaper, and not even cheaper by a factor of 2. Maybe 10% cheaper or so.
eeiei 3 days ago [-]
It’s giving desperate!
causal 2 days ago [-]
"distillation attack" is such a loaded term that really pisses me off.
Distillation is a technical term with real meaning, and historically requires logits which Anthropic does not provide.
"Generated training data" is the correct term. It's not an "attack". And Anthropic undoubtedly also generates training data for each new generation of models, yet you never see them claim Fable is a distilled Opus.
2) The word "attack" is standard security vocabulary. Per RFC 4949:
attack
1. (I) An intentional act by which an entity attempts to evade
security services and violate the security policy of a system.
That is, an actual assault on system security that derives from an
intelligent threat. (See: penetration, violation, vulnerability.)
2. (I) A method or technique used in an assault (e.g.,
masquerade). (See: blind attack, distributed attack.)
There are hundreds of named "attacks".
3) The "attack" part of "distillation attack" refers to distillers creating tens of thousands of fraudulent accounts, using proxies to bypass georestrictions, deepfaked IDs, and paying real people to pass biometric KYC checks. Who then blended this in with real user traffic to conceal their behavior.
It doesn't refer to the AI training technique in any way.
If they acquired this data without the fraud, you'd have a point.
Jcampuzano2 2 days ago [-]
Sure that can be called an attack, but then we must also concede these labs essentially massively attacked everyone else in existence to get the data, and continue attacking as we speak.
In a way you could see this as a case of Robin Hood. The US companies exfiltrated all the data on the planet just to hoard it for themselves now and accuse anyone who tries to get a piece of that back from them, and the Chinese labs are distilling it to offer it for cheap.
Obviously a bit more complicated than that but it still holds pretty well.
HarHarVeryFunny 2 days ago [-]
> Model distillation is the process of transferring knowledge from a large model to a smaller one
Sure, and large-to-small is a key part of the definition, and why it's called distillation (cf concentrating something). When Anthopic use synthetic data generated by Opus to train Sonnet or Haiku, then this can correctly be considered a type of distillation.
When Anthropic accuse Chinese companies of "distillation", it seems they are using this word to refer to two potential uses of their model outputs:
1) Using Anthropic model outputs (aka synthetic data) as training data, especially for reasoning, for Chinese models. This really isn't distillation though, since (unlike when they distill their own models) Anthropic don't actually provide the reasoning in their model output, only a "summary" designed to hide the actual reasoning. You can't distill what you are not given!
2) Another way Chinese companies may be using US LLMs is for "LLM as judge" where you are just asking the model to use it's expertise to judge/rate something that you provided yourself (to provide RL training rewards), although for coding you really want hard rewards which are easy to obtain, not fuzzy "looks good to me" ones.
Of course Anthropic are trying to pull the drawbridge up after themselves and their TOS says you can't use their models to develop anything that competes with them, and this seems to be what they are generically referring to as "distillation" - any use of their models that they suspect is being used by the Chinese to improve their own models, not just what what might more technically be called distillation, unless you want to define that word so broadly that it does mean this!
causal 2 days ago [-]
Not to mention the "distilled" models almost certainly have other training inputs as well, again watering down the meaning of the word.
And if just partial output is all that it takes to declare a model distilled, then every model trained on Internet content since 2023 is now technically a "distilled" ChatGPT and Claude model.
aqme28 2 days ago [-]
I'm not even convinced that this fits your definition. A distillation "attack" doesn't evade the security system in the sense of hacking past a login. The only part of the "security system" that it bypasses is the Terms of Service. And that's only after said data was acquired legally and correctly and normally.
It's a post-facto attack, which doesn't sit right linguistically to me.
fc417fc802 2 days ago [-]
That's understating things. The efforts to violate the ToS involve what amounts to large scale organized fraud. But I'm not convinced that should be considered an "attack" rather than merely "piracy" and I certainly don't feel like there's any ethical issue. It's nothing more than an attempt to politicize competition.
causal 2 days ago [-]
1) Your own Wikipedia link goes on to describe using logits. Yes, language evolves to mean multiple things, and that is my point: Anthropic is pushing for a watered down definition. Furthermore, Anthropic hides thinking, so you do not even really get model outputs, you get some downstream partials. Further-furthermore, Anthropic does not describe their own models as distilled when they produce training data. Why? Because generating training data != distillation.
2) This is a stretch: it allows Anthropic to arbitrarily define "attack" via TOS, and ignores the fact that the generated training data is literally paid for by the "attackers".
rarisma 2 days ago [-]
So what kind of attack are AI companies doing by scraping up copyrighted info to build these LLMs?
"You're trying to kidnap what I've rightfully stolen."
AlexandrB 2 days ago [-]
> 3) The "attack" part of "distillation attack" refers to distillers creating tens of thousands of fraudulent accounts, using proxies to bypass georestrictions, deepfaked IDs, and paying real people to pass biometric KYC checks. Who then blended this in with real user traffic to conceal their behavior.
Lol. Isn't this literally many of the same tactics OpenAI and Anthropic used to scrape the internet? So now it's an "attack", but previously it was just "training".
aroman 2 days ago [-]
Sun Tzu said: if you scrape your enemy, call it training; if your enemy scrapes you, call it an attack.
sigbottle 2 days ago [-]
Completely unrelated, but I'm seeing people and especially LLMs using causal/intervention so much it's kind of driving me insane.
It's actually a very goated term but not everything is causal, it also has precise technical meanings (although those get blurred too given that causal can mean anything from intervention proper, to mere depdnence on something prior)
causal 2 days ago [-]
Are you talking about my username? Yeah I liked the word before LLMs made it cool/uncool.
Wilsoniumite 2 days ago [-]
I like this article. Rings very true.
I like Anthropic, I don't think all their talk of safety is bluff and bluster, or at least, I want to believe that the people who left OpenAI because it had lost its focus of helping humanity still want that to be their main goal. However, yes, it seems that business fears are once again causing those in charge to turn "we want to help humanity" into "we are the only ones who can help humanity, and therefore we need to be the most profitable, and the only survivors".
If you want the former ideal to survive, at Anthropic and outside of it, you need to be willing to collaborate beyond profit incentives and recouping capex. Show other labs a commitment to research and community and they will follow. Better to bring teams together rather than implicitly say you distrust them, pushing them that way instead.
nmthornhill 2 days ago [-]
Datapoint from the cheap end of the market: I run local models on a couple of
Orange Pi SBCs and a decade-old Optiplex with no GPU. What runs usably on that
class of hardware is almost entirely Chinese open weights — Qwen's MoE builds
(35B total, ~3B active) are the only thing that gives me acceptable speed on
CPU, with Gemma as about the only western exception. I evaluated Kimi too and
ruled it out purely on size.
Whatever the strategic picture is at the top, at the bottom of the market
"weights you can download and run on hardware you already own" is the whole
ballgame, and right now that's mostly Alibaba's to lose.
chr15m 2 days ago [-]
I assume this is slow and that makes me curious - what are you using this for?
nmthornhill 2 days ago [-]
[flagged]
johndhi 2 days ago [-]
Is Gemma worse?
indeyets 2 days ago [-]
yes. less stable. more hallucinations
nmthornhill 2 days ago [-]
[flagged]
spenvo 3 days ago [-]
"Anthropic and OpenAI likely have among the lowest costs per unit of frontier-quality intelligence"
That's a big claim that his whole thesis rests on but is largely not backed up. Where are the apples-to-apples tokens-to-answer benchmarks that he's using - doesn't look like there are any, just a handwavy implication that US models are more token efficient, which they may be. But how is there so little effort in establishing this point in the article? And US labs may be in much different situations from one another: it's known that some labs like OpenAI bought big, early on compute and may have secured better pricing.
His article also does not mention the average price of electricity in China vs the US, which it seems like China leads on, and probably has the political power to more heavily subsidize. While I agree the COGS is often overlooked by top line benchmarks on coding tasks, etc, it seems that he's running on a big assumption while claiming "labs on the frontier will be fine".
c0decracker 3 days ago [-]
But.. if you are running Chinese model in the US, what difference does it make? Isn't the whole "scare" (khm khm) with Kimis is that now I don't need Claude, cause I can run Kimi on my own hardware in my own datacenter and it's maybe not as good as Claude July edition but it's is as good as Claude January edition.
richardlblair 3 days ago [-]
It doesn't need to be as good. You can route to the appropriate model and save so much money.
I have sonnet do the thinking, deepseek does all the tasks. I've massively reduced costs with this approach.
spenvo 3 days ago [-]
Sure, and I think that flexibility further undercuts his "frontier labs will be fine" take, which depends on top US labs having pricing power.
zkmon 3 days ago [-]
What's wrong if the roles of USA and China are reversed in technology? Why does the rest of the world care? It's not as if USA has done a great good for the world, and China has evil intentions towards the world. Infact it is the opposite in the case of AI so far.
over_bridge 2 days ago [-]
I feel like I'm crazy here but isn't China just sitting there doing nothing?
On one hand there's the relentless barage of American propaganda. I get that a militaristic society needs an enemy to fight against lest they turn on each other. I get that if you tell people a bad guy is coming for your jobs or your lives then you can maybe get your workers to accept worse conditions and living standards which increases profits. It's ghoulish but there's some logic.
On the other hand I can't see China doing anything except minding its own business. No tariffs. No bullying other nations. No wars started. No threatening allies. They don't let their people waste their lives on brain rot or gambling. They largely align with UN resolutions. They respect international institutions instead of always being the asterisk.
Has American propaganda just failed to work outside it's borders? It's not landing at all
jadamczyk 3 days ago [-]
I think releasing models for free is some 4d chess move by the chinese. Big chunk of the us stock market is fueled by ai mania, if the frontier labs turn out to be drastically less valuable than first believed, the downturm may be very bad. Think of all the big tech companies that have a ton of debt that they took to pour money into AI. It seems like a similar tactic to what Chinese car manufacturers are doing in Europe but the result may be more dramatic.
randyrand 16 hours ago [-]
Partial summary: By giving away weights for free, China can drive the price of white collar labor towards zero which is what the USA economy is built on. China which is built on manufacturing will not be as affected.
brasswood 2 days ago [-]
Maybe not the main point of the article, but I have a doubt about the author's introduction to commoditized markets:
> - Supplier A will sell 10 units of the commodity for $20, earning $10/unit
> - Supplier B will sell 10 units of the commodity for $20, earning $5/unit
> - Supplier C will sell 5 units of the commodity for $20, earning $0/unit
> ...
> Bankruptcy risk is where fixed costs come back to the forefront: Supplier C has both fixed costs (like potentially R&D spend) and also may have taken on debt [...] It can’t price its commodity with these costs in mind — remember, the market-clearing price approximates the marginal cost of the highest-cost unit needed to satisfy demand [...]
Why can't Supplier C price their fixed costs and debt into their product? The entire reason Suppliers A and B are earning $10 and $5 per unit, and not more, is because they cannot meet demand by themselves and are therefore at the mercy of how much Supplier C is willing to charge. Couldn't Supplier C just refuse to offer 5 units of the product at a price that would bankrupt them?
Sincerely, an interested observer of business/economics.
benrutter 3 days ago [-]
I loved this article! Regardless of how you feel about AI as an industry or tool, the economics of AI is fascinating. It's awesome to see something like this that gets into the business side a bit more.
I don't know if I agreed totally with the assessment of the risk Chinese labs pose to US labs though, in particular I think the main part I wasn't sure about was this:
> I highly doubt that Chinese models are cheaper to serve on a marginal cost basis, they just seem cheaper because Anthropic and OpenAI are so supply constrained that they are charging far more than they would if there were sufficient supply to meet the demand for intelligence.
How true is this? My understanding from Deepseek's original paper was that they focused heavily on optimising training and inference costs, in particular so that they can operate on cheaper (and more accessible to China) hardware.
It's possible I'm just not in the loop, but nobody seems to talk about US models innovating in this way (I'm just talking about cost-to-serve/train, not saying US AI companies don't innovate in other ways).
It seems to me at least, like there's a fair bit of evidence that AI shifting to a price based commodity market (vs a "best-model takes all" type market) would put China at a significant advantage? And even more significantly, require a pretty hefty correction of company valuations in the US?
mnewme 3 days ago [-]
He also forgot Europe in the equation. More and more companies use Chinese open models on European inference because of geopolitical concerns and data privacy, which could be a problem for the big US ai labs if they loose on the market.
rafmartom 3 days ago [-]
I am more afraid of the US models to be fair. A country with no clear direction in many regards , that is threatening day in and day out the rest of the World for its own interests.
ece 2 days ago [-]
I disagree with the first half quite a bit, COGS ultimately depends on the use case. If someone just wants something that a smaller model can do, running a local model on phone is going to have a negligible cost close to running any other piece of software. The alternatives to running a model also determines COGS, and even Jensen Huang has distinguished between the job and the work for the job that AI is capable of doing. Smaller models are always going to win in efficiency too.
I also heavily disagree with this no-marginal cost in software distribution view whenever I see it, bit rot is real, and someone is paying a marginal cost whenever they do an update. You have to re-distribute with changes whenever anything changes. These costs are just hidden because things are ad-supported or bundled in some way. These costs are also kept low because of standards and open source, but could become high anytime. Additional licensing also has costs.
That said, I couldn't agree more with the last paragraph, charging a high price for models would be better than denying access for any model that wants to stay relevant.
simonreiff 3 days ago [-]
I fully agree with everything in this essay. Make distillation fair use. And let us use Mythos/Fable and Sol and successor or future models for all cybersecurity purposes.
jdw64 3 days ago [-]
While intelligence is said to be a replaceable commodity, oil and copper can be used in nearly the same way even if you change suppliers as long as the quality grade is matched. However, I question whether two models that produce the same benchmark answers are actually interchangeable in real world use.
Personally, I think models will increasingly become specialized in different areas, some good at X, others good at Y, and we might see workflows that mix multiple models.
vfalbor 2 days ago [-]
You couldn't be more wrong. What's being sold are chunks of time with access to specialized hardware resources. Through which model or with which device is irrelevant; the one winning, and continuing to win for some time, is Nvidia, and that company is American. As long as they have the H100, B200, or B300, China won't be able to compete with the American strategy, no matter how many new models they release, because these types of cards require incredibly powerful hardware to run.
oezi 3 days ago [-]
The article makes a great point that the token industry is going to be commoditized as time goes on.
Following this argument the key for each player will be the underlying cost structure and serving capacity to offset the upfront R&D cost.
The cost infrastructure will be driven by access to cheap electricity and cheap chips. The capacity will be driven primarily by depth of pockets now to buy all available supply in chips/mem/data center building capacity. While China is certainly in the lead on cheap energy, I am wondering if they can/want to beat the > 1tn USD being spent on data centers right now. Following the example in the article:
If company C from China sells 10 units for 20 USD produced for 10 USD they pocket 100 USD.
If company A from America can sell 100 units for 20 USD produced for 15 units, they pocket 500 USD or 5/6th of the market's profits.
softwaredoug 3 days ago [-]
> By the same token, don’t expect China to do anything about distillation attacks on the frontier labs. I think it is mistaken to attribute all of the success of Chinese labs to distillation, but it’s just as much of a mistake to pretend like distillation doesn’t give Chinese labs a big advantage.
I think we see this with Meta being paranoid about internal Claude usage, to avoid inadvertently distilling[1].
If distillation is a driver, then smaller American labs could be distilling, but are not for legal reasons.
No one should be afraid of anything. Fear is a terrible advisor. Keep your eyes open, try to read the context as careful as you can and adapt as best as you can. Don’t spent too much time trying to be an oracle, never works out…
hexator 3 days ago [-]
I'm worried that any ban on Chinese AI models might be an excuse to get mass surveillance.
onesociety2022 3 days ago [-]
You don't need mass surveillance to enforce such a ban. Once the US Govt declares Chinese AI models are banned, no US business will use them nor distribute them. Any cloud service that rents out GPUs in the USA will explicitly prohibit the use of Chinese open model weights in their terms of service (you open yourself to a lawsuit if you violate their ToS). Any Tokens-as-a-Service provider will refuse to serve those tokens to customers in the US.
Sure as an indie hacker, you could go download the weights for a Chinese model with a VPN, and then attempt to run it at home by building your own GPU cluster but these large models require quite expensive hardware to run on and so it makes it less likely than anyone would invest that much capital to do something that is illegal. There's no way for them to sell a legal service using those tokens. So it can only be strictly for personal use (the Govt won't care because very few people will have that kind of money and risk appetite). The other option will be that there will be some shady third-party providers in foreign countries who are willing to sell tokens from these models to US consumers knowingly.
chockablock 3 days ago [-]
> Any Tokens-as-a-Service provider will refuse to serve those tokens to customers in the US.
So under a ban rest-of-world gets to use cheap open-weight models but American companies/individuals must only use only ‘approved models from US for-profits’? Doesn’t seem like that kind of protectionism will be popular or politically tenable. Not so long ago US chose cheap TVs over maintaining the country’s manufacturing base.
(Despite what you wrote it’s also really hard to imagine that enforcement wouldn’t leak like a sieve. Unser sufficient economic incentives [which are the predicate for the ban], loopholes will be found.)
softwaredoug 2 days ago [-]
Something I did not realize. The White House instructed intelligence agencies to help in preventing distillation of US models.
Whether or not distillation matters a small amount or a big amount, still interesting:
The author doesn't seem to realize that a healthy margin has been built into the inference pricing. Once low cost open source inference providers get their hands on powerful frontier level models, there would be a severe margin compression for OpenAI and Anthropic.
Why do you think an inference provider competing for the same compute as OpenAI and Anthropic gives up those margins rather than giving a modest discount over the frontier for near-frontier performance?
mattas 3 days ago [-]
"Right now, none of the above analysis applies because demand exceeds supply for frontier models, and supply is limited by a lack of compute."
It gets particularly hairy because models themselves can tune their "token verbosity" to manufacture demand for compute. If compute was such a precious resource, you'd think we'd be complaining that the output was too terse.
The ability for a vendor to determine ex post facto how much a query costs is a similarly new economic phenomenon to zero marginal cost.
ayunuse 1 days ago [-]
I think chinese componies are not doing charity for releasing their models publicly. That is the strategically best decision they can do for now. Morally they should but IMHO they are not angels :D.
golly_ned 3 days ago [-]
> I expect the inference market to grow much faster than training costs
This was my assumption as well. It's also generally true of 'traditional' deep learning models that inference cost is expensive compared to training.
But the cost per token for inference has been very quickly dropping. I don't recall where, but I recall about ~50x down from GPT3, even as model complexity has increased. Even with agentic systems, there are lots of optimization opportunities. I'm less assured about claims like this.
overfeed 3 days ago [-]
> [Anthropic/OpenAI] are serving models at a particular capability level for months before their competitors, and are simultaneously applying the best models to optimizing those costs. Second, intelligence isn’t in fact a perfect commodity, in part because applied intelligence makes itself smarter
Is he casually assuming a singularity has already happened? A regular first-mover advantage I can understand, but those have been squandered or lost many times before.
magarnicle 3 days ago [-]
No, he was just talking about reinforcement learning, etc.
overfeed 2 days ago [-]
So that's just an appeal to authority (longevity?) not anchored in reality. Just because you've been doing something for longer doesn't mean you're the best at it; Google has been shipping AI/ML models years before the founding of OpenAI and Anthropic, but it's playing catch-up on LLMs.
softwaredoug 3 days ago [-]
Haven’t we been in this “China is 3-6 months behind” for a while now (maybe up to a year? Longer?)
The actual difference is how much scrutiny and time was put into the Mythos / Fable and GPT 5.6 release. Making it feel like “these are a big deal”. Spring and summer THAT was the AI story
Then Chinese labs release models that approach Fable performance. We’re shocked they just seemed to appear out of nowhere.
It’s less about the gap closing. It’s more about the weight we put into Fable-capable models.
podgorniy 3 days ago [-]
Just couple years ago mr altman was promising open AI, to benefit humanity. They even forgot to remove this part about "openess" from the company name.
Today chineese deliver that promise and usa people freak out like they have any skin in this game. Enjoy the ride leader of the free world....
olalonde 3 days ago [-]
I still can’t wrap my head around how transferring all assets from a non-profit foundation to a for-profit company was legal.
kiicia 3 days ago [-]
altman should listen to his own words, and he should replace himself (wasting energy for eating) with ai (supposedly not wasting energy for eating)
it would be beneficial both for openai and world (most likely)
minraws 3 days ago [-]
Me I am, so very afraid of actually decently priced inference.
kautryii 2 days ago [-]
I'm afraid of Chinese models because they are rip-offs of other models...and there's no telling what your data is being used for when you utilize the API. It's one thing to have US companies using my data and another to have a Chinese company who aren't bound by any IP laws jacking all of my codebase.
desterothx 2 days ago [-]
IP laws work if the penalty is larger than the use case of the data. Pretty sure literally every US company is using the data knowing they won't be punished equivalently. We're living in the age of trillionaires, Facebook paid 5b$ for the cambridge analytica scandal, something Elmo could write off as a business expense at this point...
causal 2 days ago [-]
Open weights dude. You can literally run it on your own or rented hardware and give your data to exactly nobody, unlike closed models.
lenerdenator 3 days ago [-]
It really is amazing that China went from the country that hacked Google out of its market to a trusted source of AI in the tech world.
Joel_Mckay 3 days ago [-]
Distilling models using more advanced LLM is not a new phenomena. It is a cost effective strategy in a highly competitive emerging field.
Also, the Hidden-Agent problem exists in every model, and is a persistent tangible risk independent of whatever team people cheer for at the games. Let us remember, every LLM nuked all of humanity 92% of the time in simulated war games. =3
samiv 2 days ago [-]
I'm afraid to trust technology coming from companies operating under the jurisdiction of a rogue aggressive nation that is continuously attacking other nations both (so called) ally and foe using economic and military actions.
flybarrel 2 days ago [-]
Sir, given what's happening in the world these days, I cannot really tell which nation you meant by your statement :D
joering2 2 days ago [-]
Really? Ok, name list of rogue aggressive nations that are continuously attacking other nations both allies and foes, and I tell you which one he meant...
rnd0 2 days ago [-]
Are you including aggressive economic attacks -such as weaponizing tariffs or are you only counting direct military acts?
I'm pretty sure the U.S. has done both over the last couple of years -but I'd love to be proven wrong. :)
drop_star 2 days ago [-]
The United States
joering2 2 days ago [-]
its hard to chose from your list...
DarkNova6 2 days ago [-]
Well done, sir. This post has evoked the expected responses from the other comments.
bigyabai 2 days ago [-]
Are you describing the United States or China with this quote? It's hard to tell.
tcmart14 2 days ago [-]
Outside of a few boarder disputes with India, I don't China has militarily attacked anyone since they got their ass handed to them by Vietnam (Sino-Vietnamese War 1979). So I think that rules out China.
stronglikedan 2 days ago [-]
Hong Kong would like a word...
therion93 2 days ago [-]
Hong Kong is a Chinese city since 1997. The aggressors in this whole situation are the British bastards
smitty1110 2 days ago [-]
The Philippines have some news for you…
2 days ago [-]
therion93 2 days ago [-]
Are you talking about US companies, right?
stronglikedan 2 days ago [-]
your loss ::shrug::
gigatexal 2 days ago [-]
Not saying the heavy hand of the Xi admin and Chinese communism is equivalent to the outright lawless, corrupt, grifting and vindictive and hateful administration that is Trump 2.0 ... but it's pretty much * makes the 6 7 motion that the kids do * this for me.
China will just do a better job -- if they do this at all -- of storing, collating, indexing, and using data from their state sponsored and championed AI labs to use that against the US. [1]
The US under this admin is doing the same, attacking universities, allies, it's own citizens.
The two governments are operating more or less the same. Ergo the ai models from each country's ai companies shant be trusted either.
A reminder any comment about risk FROM china, invites a "Tu Qoque" facing the other way. The paranoia here is probably fully symmetrical.
I see massive risks in belief the inferences drawn from strategic information cannot be seen. So if you depend on some position remaining inside a secure facility but you drove to it from data outside that secure facilty, The likelihood that an inference model can derive the same idea is very high. Collation over public data is not inherently secret because you used a secret model or secret weights.
A more simplistic take might be that the fear is not actually driven in the secrets, the fear is "the emperor has no clothes"
JimsonYang 2 days ago [-]
Why hasnt the EU develop an competitive open source model? Ive only of mistral, but with the chinese models you have GLM, Kimi, Qwen, Deepseek etc all of which seems to be better and better
random3 2 days ago [-]
Because the European researchers and engineers are busy building models in the US
JimsonYang 2 days ago [-]
Is there more info?
3 days ago [-]
holoduke 3 days ago [-]
So happy that we have finally 2 countries playing the competitive game. No more secret deals between competitors. No real competition. A race to the bottom is always a good thing for consumers.
ArtRichards 3 days ago [-]
New, smaller models can outperform the previous generation's foundation models.
What if there's a way to extract the commodity of intelligence from smaller models?
I've seen for many use cases it's well enough. :)
zuzululu 3 days ago [-]
My thinking is that with the current narratives out of washington we are on track for a ban on Chinese models and possibly sanctions against Chinese AI companies
I think it is the right move to protect American interests
ChrisArchitect 3 days ago [-]
Related:
Ben Thompson is wrong: US frontier labs are right to be panicking
They are just afraid open source/weights models, and the models are from China.
If EU build some SOTA open source models, they will design a different story
ilamont 3 days ago [-]
But it’s a problem to be dependent on China. The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation, for U.S. companies at a minimum.
I'm amazed that no one is talking about proposals that are surely being discussed in Washington and pushed by SV lobbyists to restrict Chinese models on national security grounds, or other some other basis.
The belief that Bytedance could engineer a finger on the algorithmic scales to serve the interests of the Chinese Communist Party led to a lot of debate in Washington, and ultimately resulted in TikTok being divested from its Chinese owners. Huawei is shut out from the U.S. market, which limits its business even in markets where it's not banned because it's effectively stamped with a scarlet letter.
IMHO, Chinese models are headed for a similar fate or at least a showdown in Washington or the courts because they are supported and/or controlled by entities which ultimately serve the CCP.
caruasdo 2 days ago [-]
US restrictions are excessive, and companies have to protect their infrastructure with the same technology they are trying to block.
zzzeek 3 days ago [-]
the leader of China praised Open Source in a speech. Crazy times
nl 3 days ago [-]
> because U.S. open weight model makers must follow the frontier labs’ terms of service, they (1) are worse than Chinese alternatives and (2) end up distilling the distillation, just with a detour through Chinese labs. Wouldn’t it be better if western open weight model makers could go to the source?
Is this an assertion that is backed by evidence?
From the Elon/OpenAI trial:
> On the stand in a California federal court on Thursday, Elon Musk was asked if xAI has used distillation techniques on OpenAI models to train Grok, and he asserted it was a general practice among AI companies. Asked if that meant “yes,” he said, “Partly.”
The best model is the model that runs best on your hardware.
sperr11 2 days ago [-]
I think a simple experiment is enough to understand why one would have some concern with a state-censored AI model : Just try asking them about atrocities committed by their state[1].
It's hard to overstate how important this point is. If China is the standard of openness, it's a pretty low standard. Your point I think adds to what the article is saying, from a different perspective, but arrives at a similar place. Our best selves in many ways are defined by openness, unflinching self reflection, and competition. We should remember that, and as the author of the article says, lean into it rather than letting fear mongering and histrionics protect these models from competition.
rnd0 2 days ago [-]
Yes and no? I've had the same experience with asking about Tienanmen square -but then when asking a Deepseek v4 model (and confirmed by asking the chat on the deepseek site) about the "laying flat" movement it gave me a detailed answer that was unexpectedly sympathetic to the movement.
YetAnotherNick 2 days ago [-]
ChatGPT supports left wing American narrative and according to them Vietnam war was American's fault. Could you try asking people killed due to left wing American ideology.
tombert 2 days ago [-]
ChatGPT tends to support any narrative that it thinks the user supports.
I just asked ChatGPT about the Vietnam war and it did not say that it was purely the US's fault: https://imgur.com/a/zmiOyuu
It also didn't seem to have a problem describing people killed for left wing ideology: https://imgur.com/a/LhH9saL
These are both with the free ChatGPT membership, as I do not have a paid membership anymore.
I know this is a common trope to bitch about, but honest question: did you actually try this before you commented?
ETA:
I was curious what something that was trained around me specifically (fairly typical lefty American progressive) would say, so I asked Claude (which I have a paid membership for and have discussed political things about many times). The answers were broadly similar: https://imgur.com/a/caFgKHH
2 days ago [-]
bandofthehawk 2 days ago [-]
What do you mean by America's fault? Maybe I'm in a left wing bubble, but I was under the impression that it's generally accepted by left and right that the Vietnam war was not justified and did not accomplish its stated goals.
samtheprogram 2 days ago [-]
Abortion? Permissive drug laws? Immigration? Homelessness policies? Vaccines? American AI gave me these examples, and they're not really so cut and dry. Do you have a better example that's being left out (no pun intended)?
johndhi 2 days ago [-]
Good Lord I hadn't ever heard of the mai lai massacre. Just spend 10 mins reading about it. Christ.
MetroWind 2 days ago [-]
The real scary part is China's compute capacity is orders of magnitude smaller than US's.
alizaki 3 days ago [-]
There is no “Chinese LLM”. Each “lab” is distinct and their models behavior is as unique as those from OpenAI and Anthropic
wmf 3 days ago [-]
Somehow a certain set of labs are all releasing open weights and a certain other set of labs are closed weights.
dofm 3 days ago [-]
Somehow the two main closed weights frontier models come from two companies with HQs about two miles apart, and the CEO of one used to work for the other.
gagandeeprangi1 3 days ago [-]
Thinking machines will do it for usa
NooneAtAll3 3 days ago [-]
I don't understand the premise in the beginning
how is running servers supposed to be 0 cost, while running ai inferrence isn't?
cheema33 3 days ago [-]
> how is running servers supposed to be 0 cost, while running ai inferrence isn't?
For a SaaS business, running servers isn't free. But compared to the cost of running GPUs for inference that you are selling, it almost is. The company I work for is a SaaS company. We have a single production server. A couple of QA servers. All hosted on Hetzner. Monthly cost for servers is less than $400. This generates a few million dollars a year in revenue.
If we were in the business of selling inference, our cost of providing the service, for the same amount of revenue would significantly higher.
Even large businesses like Microsoft, Meta, Google have operated with similar margins. Cost of running servers, compared to revenue was very low. But inference changed that, in a dramatic way.
throwawayffffas 3 days ago [-]
A typical server that costs 10k to 30k to own and operate can serve between hundreds and thousands of requests per second of a traditional web application like facebook for 2-4 kW of power, the marginal cost of each request is effectively zero.
A single response from kimi k3 requires hardware that cost between 500k and 1m dollars up front and draw over 20kW. Each request costs at least 5% to 10% of the charged cost.
kmeisthax 3 days ago [-]
> To that end, here’s an even more interesting question around distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here?
Frontier labs that thought they could Rupoor[0] the entire creative class, transferring the coercion premium of copyright ownership from Hollywood to themselves. In their eyes, copyright should not apply to them, but also their models should have exactly the same value as a copyrighted work.
Stratechery also argues the US should explicitly make training fair use and forbid terms of service that prohibit distillation. I'm in support of the latter, but NOT the former, even though I normally hate copyright. My reasoning is primarily that copyright is one of the few legal paths available for a rando to go and put the work of an AI frontier lab in legal jeopardy. In the EU and Japan, such legal action has already been foreclosed by similar law. And while free distillation would obviously be preferable, it's also much more of a legal long-shot. Getting America to do anything that even smells like taking property away from the powerful is impossible[1] - it's our zeroth amendment. But we can at least hack the property laws that currently exist to cause problems for the frontier labs.
And, to be clear, if distillation is OK but training is not fair use, distillation is still OK. The output of an AI model is never copyrightable, because copyright only protects the human element. Essentially, this would say "don't train on humans, but absolutely rip off and steal the shit out of other AI labs and give it to the rest of us."
[0] In the Legend of Zelda series, Rupoor is anti-money - collecting it decreases the amount of money you own. I am using it to mean "turn someone's asset into a liability".
[1] Given that America was literally created to protect a wealthy land/slave owner class from disenfranchisement, either from above or below, and the last time we did this we literally had to fight a civil war against that same owner class that installed a new owner class that has largely remained today
chews 3 days ago [-]
later secondaries investors in openai/anthropic. It's like time traveling into the spacex ipo.
Havoc 3 days ago [-]
oh wow - hadn't realized they decided to opensource Qwen 3.8 Max. That's pretty big news.
tesnorindian 3 days ago [-]
Are the open weights models accelerating layoffs in Indian IT sector?
chimp_brain 3 days ago [-]
They are neither accelerating or decelerating layoffs in Indian IT sector, from what I know they were bound to happen as software and IT were never hard skills. They were anyway supposed to move to Africa or other countries within a decade, so a lot of executives were cautious even before 2023. AI has simply killed the opportunity for an Infosys or TCS in countries in Africa or even other South Asian countries.
Der_Einzige 3 days ago [-]
The world would be so much better off if the WITCH companies ceased to exist
tesnorindian 2 days ago [-]
This will create a massive entrepreneur boom in India due to low cost access to AI. Hope these entrepreneurs don't become witch with black magic.
tesnorindian 3 days ago [-]
Interesting and true perspective.
ab_wahab01 3 days ago [-]
Honestly, as someone from a developing country, this shift is good for us. US frontier models are too expensive for us to use regularly. Chinese open-source models/subscriptions are really good to use.
throwitaway222 3 days ago [-]
A company making a decision to allow use of chinese models is a company also choosing to send tons of various credentials to chinese model companies. These will just get scooped up, OpenAI and Anthropic can probably hack into anything at this point if they wanted to.
Tostino 3 days ago [-]
But, they are releasing the weights very shortly (or already have for some of the models discussed). For a very large company, you can purchase or rent the hardware yourself to serve the models.
Or any US hyperscaler with GPUs to spare can decide to serve the models for a reasonable cost/token.
You don't have to send China your data.
purplepatrick 3 days ago [-]
Commenting wholesale on some folks who are asking for hard evidence. I cannot provide that either but can contribute some empirical data.
I have been working on a project with about a dozen generation tasks, each of which comes with a fixed token budget. The nature of this system requires that most tasks be completed by distinct model families.
As a result, I tested ~50 models across as many model families as I could gather, frontier and open weight, API (gateway and direct) and self-hosted. Evaluation was based on a set of cosine similarity validations that was repeated across ~50 different embedding models.
Interestingly, frontier models did worse on the tasks than open weight models. However, when it came to costs, the picture was reversed: frontier models were much, much more token-efficient. In fact, almost no open-weight model was able to meet the initial token budget, while almost all frontier models did. Moreover, open weight models struggled massively with reasoning, in terms of latency and token consumption.
I also found that the latest models did not perform better than older models. And any a priori benchmarking data was utterly useless.
So, I ended up using a set of open weight models without reasoning, as it turned out reasoning as well as frontier negatively correlated with the tasks. However, before I knew this, I had spent a lot of time running each available reasoning level for each model.
Lastly, as an aside, when it came to embedding models, size (dims as well as model size) did not correlate with quality, once a hurdle figure (~2k dims) was met. In fact, sweet spot was 3-5K, and for my (text-based) set of tasks, dense models tended to outperform MoE ones.
josht 3 days ago [-]
Someone (anyone!) get David Sacks on the horn and tell him to read this.
Talpur1 3 days ago [-]
I personally believe opensource or may be state owned LLMs are future, every country on earth should have its own national LLM,trained on country's own data, and then allow its public to use it for free
sva_ 3 days ago [-]
I see a lot of ways how that could lead countries into a dystopian nightmare, where the gov gets to decide what kind of biases a model should have.
jonathanstrange 3 days ago [-]
Gemini Pro has become so bad for my purposes -- copy & paste Go programming and code analysis with the web GUI -- that I'll take any model with equivalent capabilities at the same price or lower. I don't care where it comes from, I'm not dealing with state secrets and, frankly speaking, US corporations have an abysmal track record regarding safety and surveillance.
55555 2 days ago [-]
There are really only two factors at play here: people trying to protect their massive investments, and governments fighting over who gets backdoor access to all of your chats.
nunez 3 days ago [-]
I really enjoyed reading this.
This might be a simplistic take, but my biggest worry with depending on Chinese models (and, by proxy, open-weights model development) is that the US can deem them a national security risk at basically any time, and Ant/OAI have minimal interest in making frontier-level models open-weights.
Regulated companies prohibit Chinese models in anticipation of the ban-hammer from the feds, so for data-sensitive work, they're stuck with LLaMa, gpt-oss and Gemma models (which are good and serve as a good-enough base for sft, but seemingly not as good or as expensive as Chinese models)
I suppose the USG can do the same thing that China is doing and bankroll/subsidize that effort; whether they will is for fate to decide.
Nonetheless, this article made it clear that nVIDIA is the real winner in all of this. Shovel selling to the extreme.
pupskipper 3 days ago [-]
The fact that Anthropic has a model like Mythos means that counterpart countries like Russia and China are not far behind, if they haven't already developed something similar or better.
konart 3 days ago [-]
As a Russian: no. Sber's GigaChat 3.5 (released two weeks ago) is the best we can right now and it's years behind SOTA models.
3 days ago [-]
anuramat 3 days ago [-]
> Russia
lmao
iLoveOncall 3 days ago [-]
All the article relies on the premise that selling tokens is profitable. I don't see any indication for this, and it makes the whole house of card crumble.
nottorp 3 days ago [-]
It's say Anthropic, Allegedly OpenAI...
mghackerlady 2 days ago [-]
The american AI companies, presumably
jdla1o 3 days ago [-]
The future is SLMs and China is going there..
panchtatvam 3 days ago [-]
Is this article written by AI ? It looks so.
jp0001 2 days ago [-]
American models == Cars with no brakes.
Chinese models == Cars with brakes that can't drive down tiananmen square.
phkahler 2 days ago [-]
>> Going forward, however, I expect the inference market to grow much faster than training costs (and that includes the assumption that training costs will continue to skyrocket), which means they really can make it up in volume.
But as inference becomes cheaper, some of the market will move to self hosted inference. I look forward to someone supplying small servers designed to run inference locally.
With or even without open models these companies are selling compute, and we've been making that rent vs buy decision for 60 years.
daitangio 3 days ago [-]
I think the most interesting part of the article is the Huggingface incident at the end:
>Right now defenders are effectively banned from using Fable or Sol for cybersecurity because of Trump administration directives; that means the best alternative is using models from a country which has been trying to weaken our cyber defenses for years. This is insane!
I understand some guardrails are needed, but it is becoming increasing problematic manage them without a strong public discussion.
2 days ago [-]
adrienfr31 3 days ago [-]
I am French, and I am sad to see that French models aren't being talked about and that it remains a battle between the Americans and the Chinese.
alfiedotwtf 3 days ago [-]
Answer: US investors
namar0x0309 3 days ago [-]
I haven't had the time to look into recent and past history, but my intuition points at failed empires having similar "elites" starting to stagnate innovation for the sake of "protectionism" - whatever that means.
vachina 3 days ago [-]
Nothing changed for China. The only difference between then and now is that now China is the one selling the products, instead of western capitalists taking a cut off the COGS and selling price.
You techbros need to get off your ass and go to work.
rib3ye 3 days ago [-]
Why hasn't someone created a disinformation benchmark to assess that part of the argument?
Alien1Being 3 days ago [-]
The options are to use LLMs from a country run by a psychopathic regime or alternatively to use Chinese LLMs.
marwaneet 3 days ago [-]
i think most is vcs
icase 3 days ago [-]
not enough people
kdqed 3 days ago [-]
I'm honestly more afraid of Claude
npn 3 days ago [-]
what a horrible article. full of misinformation and dishonesty.
1. training new base models are expensive for sure, but fine-tuning them are relatively inexpensive enough the labs can continue to do so forever. the main reason why frontier models are so good is because the massive input they generated from user usage. they are using that information to strategically build better training data. and this is why no other models can catch up, til now that is.
but if chinese models are good enough, and free to host, and cheaper to use, then the consequence is the frontier labs will lost valuable user inputs and the chinese labs will gain more. as time goes by this will be a domino effect.
2. nvidia is not only the player in the hardware scene. amd mi350p is getting popular, and huawei is pumping SuperPoDs. what does this mean for us? chinese models will surely use chinese hardware, and optimize for them. the other people will pick amd because compare to nvidia they are cheaper. with open weight models and open source inference stacks, they are freely to experiment and improve the stack, thus further lower the inference cost and nvidia dependency.
and they even plan to build their own inference hardware, too.
and nvidia loses market share meaning all the fund it gives to openai or anthropic will be cut, too.
and you say there is nothing to afraid?
oska 3 days ago [-]
> Second, intelligence isn’t in fact a perfect commodity
We have got very far from Cicero's coining of the word 'intelligentia' (from inter legere, a 'reading between' and hence discernment) when people talk about 'intelligence' as a commodity
People have been decrying the 'cheapening' of the word intelligence for over a century now, going back to Psychology's adoption of the word and coining of nonsenses like "Intelligence Quotient". "Artificial Intelligence" is just the latest degradation of the original humanistic meaning, and now people aren't ever bothering to prepend 'artificial' to their idiotic use of the word
veloxxn 2 days ago [-]
[flagged]
troygentic 3 days ago [-]
[flagged]
supportm 2 days ago [-]
[flagged]
hereme888 2 days ago [-]
[flagged]
nttylock 3 days ago [-]
[flagged]
andrewdubinsky 3 days ago [-]
[dead]
ngl999 3 days ago [-]
[dead]
YouDontKnowAI 2 days ago [-]
[dead]
yoshitha2010 2 days ago [-]
[flagged]
rahul1891 1 days ago [-]
[dead]
martinbfine 2 days ago [-]
[dead]
stiltzkin 2 days ago [-]
[dead]
threerouter 3 days ago [-]
[dead]
brandopn 3 days ago [-]
[dead]
warshinder 3 days ago [-]
People with 401k’s and retirees?
nnm 3 days ago [-]
Looks like authored by AI. With quite some reasoning, but no real data to back up main point.
sharadov 3 days ago [-]
What makes the Chinese models this good? I don't believe it's distillation alone.
This from OpenAi's Head of Strategic Futures
"Some observations on Kimi:
It's a very good model! I don't think its performance can be explained away by distillation or anything like that"
China's strategy of spending billions on training these models and open sourcing these models away is strategic - they want to kill the US LLM industry at any cost.
To win on the AI front by any means necessary.
nl 3 days ago [-]
> China's strategy of spending billions on training these models and open sourcing these models away is strategic - they want to kill the US LLM industry at any cost.
Why is it when Anthropic and OpenAI spend billions trying to beat each other it is competition, but when the Chinese companies do it then it is trying to kill the US LLM industry at any cost.
The US federal government spends billions in subsidies via the US Chip Act, and bans chip sales to China to support US companies.
But the implication is that somehow Chinese competition is illegitimate because "strategic".
rightbyte 3 days ago [-]
Jingoism is the answer I believe.
sharadov 2 days ago [-]
I am not defending the US LLM industry or the government, all I am saying is it's similar to an arms race.
I did not suggest anywhere that what China is doing is illegitimate.
China has state supported capitalism, and they will do whatever it takes to prop up the AI industry. Anthropic and OpenAI want the same protections.
Havoc 3 days ago [-]
>What makes the Chinese models this good?
Why wouldn't it be? China is pumping out AI research and researchers at a staggering pace and there is no inherent reason why western models should be better
oeatwell 3 days ago [-]
I think frontier labs should start building ecosystems by partnering with companies that already have software products, and even collaborating with hardware manufacturers. The ultimate goal should be to create a much broader range of products that integrate naturally into people's everyday lives.
The United States' real advantage over China is freedom. Chinese LLMs simply can't compete with American ones when it comes to the humanities, creativity, entertainment, or financial transparency. As long as the U.S. continues monetizing these strengths, the compounding effect will make it virtually impossible for China to surpass the U.S. at the product level.
samtp 3 days ago [-]
This is a really ironic comment given how the US gov is getting politically involved in pretty much all science research funding and speech at the moment.
oeatwell 3 days ago [-]
Exactly. If this continues, the US will be fighting the wrong competition against China instead of building around its own structural strengths.
Aperocky 3 days ago [-]
> humanities, creativity, entertainment, or financial transparency
> monetizing these strengths
> real advantage over China is freedom
Please tell me if I'm unfairly paraphrasing but these seem to be your main argument and they seem to be oxymorons
oeatwell 3 days ago [-]
[dead]
sjreese 3 days ago [-]
Kellogg School of Business -- he said -- token as a commodity and therefore Open AI is constrained .. ha ha ha hee hee ha .. Well... you build a better mousetrap, and DeepSeek, K3, and ByteDance are just that -- just as good and fit to purpose -- What is needed is to build on top of -- not paniteir (invade privacy and kill people with the information) -- not USMC AI -- use PI's as overwatch killer drones -- but how can I make harder steel, longer-lasting, seawater-resistant concrete, faster time to build housing, better enforcement of USDA rules and FDA adverse enforcement, and better EPA water cleanup, a better FTC for consumer goods -- that is, if I buy an item, that item is safe and built to purpose -- ANYONE not talking about public protection of consumer rights usng AI, is wasting your time
magarnicle 3 days ago [-]
Has someone replaced your return key with a double-dash key?
sjreese 2 days ago [-]
ok look at this way, without being petty what is your ability to have clean water today? and how is that measured? -- You know the food is substandard to the USDA standard EG ( Taylor farms 2026 ) But Where is the enforcement for Your drinking water and the food YOU eat daily. And what is the projected outcome in years from eating substandard foods to NIH standards? -- All of this is known today. And AI can help, by YOU building on top of the models you have access to in 2026. The idea that tokens are constrained - is the wrong approach in Business and public policy. You may recall how KSG got its funding and why? In todays world YOU have a responsibility to build better - As such a better mousetrap.
1) China can (and does) use the models to influence the west. They train in false information about Taiwan and Hong Kong. Or pretend like history is in favor of China.
2) Ignoring the models containing false information, they are incredible. But you should be scared of running inference via the model creators directly. If you think your data is safe compared to running it via model providers in the US ( either frontier or model hosts like fireworks.ai ) then please let me know your bank details so I can poke around.
Why should I trust a US company more than a Chinese one?
I mean lesser of two evils thinking, if one is intentionally leading us towards climate disaster, while the other isn't then yeah. What else can be said? Should I trust the authoritarian country who believes in engineering and science, or the one that doesn't?
China brought 80GW of new coal power online last year. The US added 0, and plans to add 3GW next year (we all know why).
China doesn't care about climate change, they care about energy independence, and conveniently have very little natural fossil fuels besides coal. Which they heavily mine and utilize.
Although it is highly suspicious why they are building so much unused capacity, it is as if they expect something to happen soon.
I certainly agree there's other motives (renewables even net them significant trade, like EVs) but it's not as though 'caring about' climate impacts requires selfless virtue, they stand to lose $trillions/year like everyone else.
Geolocated here: https://www.google.com/maps/place/24%C2%B024'57.0%22N+113%C2...
I asked ChatGPT how many coal train cars are required per year to fuel this 4GW plant: About 200 coal train cars per day. That is staggering! It is about 5.25 million tons of coal per year. Multiple that by twenty and that only covers new capacity added last year. It is depressing to see these numbers.
Are you sending them extra money to fund nuclear?
https://www.worldometers.info/co2-emissions/co2-emissions-by...
The US emits 49% more CO2 per capita than China. And even with the much larger rate of increase, 0.79% to US' 0.3%, it'll still require 81 years to catch up to the US per capita rate [0].
The US aren't the good guys here, not by any means. Compared to my country, for instance, the US' emissions are 3x per capita that of the UK. If you want to argue that we shouldn't be using per capita (although we absolutely should), then that's 45x the UK's CO2 emissions.
[0] log(13.59/9.13) / log(1.0049) = 81.4 years
That is not the actions of a country who “believes in climate change”.
China’s CO2 per capita is ahead of every large developed country except for the US, Australia, and Russia.
Because those are the only ones that can move the needle on climate change.
No one cares if Iceland emits more co2 per capita. They have < 400k people.
This isn’t about fairness, this is about whether the actions of the leaders of China care about impacting climate change.
https://ourworldindata.org/grapher/imported-or-exported-co-e...
US numbers are insanely high because cheap hydrocarbons are locally available (=> bad incentives) and everyone is wealthy (that correlation is very strong; just compare Luxembourg, which is much wealthier and more polluting than surrounding nations) and also population density is rather low so more energy wasted for transport.
Well, your theory does not hold at all if you look at Switzerland which pollutes 1/4 of the US per capita. CO2 pollution is not and does not have to be correlated with standard of living.
If you want meaningful CO2/capita comparisons, you also have to be very careful with countries that get ("free") hydro power opportunities for electricity because that distorts the picture massively (same for e.g. Norway).
Big producers of hydrocarbons, on the other hand (like the US) have to work hard to resist the allure of cheap & convenient fossils.
Switzerland is so far ahead in this comparison because they get a lot of CO2-free hydroelectricity (>50%), have to import most hydrocarbons (=> incentive against) and also save massively on transportation because density is much higher.
Some fraction certainly is. Here is trade-adjusted CO2 per capita: https://ourworldindata.org/grapher/consumption-co2-per-capit...
Note how this still trends down for most western industrialized countries despite positive economic growth over the last two decades.
Consider the gradual takeover of SUVs on the roads (many which to be fair are probably about as efficient as sedans were 20 years ago, albeit more expensive in real terms and significantly more dangerous for pedestrians), or the heavy reliance on trucking rather than rail for heavy freight transportation between logistics hubs, or the heavy reliance on air for long-distance travel due to having literally zero high-speed rail, or the fear of nuclear power, or the overt political opposition to green energy.
Just in comparison to Europe “46 percent of European freight goes by truck while only 11 percent goes by rail, while in the United States more than 40 percent goes by rail while just 30 percent goes on the highway."
This is in ton-miles, so it’s the metric most directly relevant to emissions.
According to eurostat (https://ec.europa.eu/eurostat/statistics-explained/SEPDF/cac...) that is two thirds of freight volume and it probably still eclipses even rail transport in CO2 efficiency (until full CO2 free rail electrification), so this probably distorts things completely...
If you actually compare the emission fractions, you can see that the US uses much more CO2 per capita on transport compared to Europe (about 3 times as much!):
https://ourworldindata.org/grapher/per-capita-ghg-sector?cou...
This report from Eurostat is where I got my EU numbers.
https://ec.europa.eu/eurostat/documents/3217494/5711595/KS-D...
Maritime is around 37% if you only include domestic and intra-EU shipping.
Your report include all shipping that passes through any EU country’s EEZ. So if a ship coming from Africa headed to the US passes by the Azores. A portion of the tonne-kilometers of that ships voyage will show up in your report.
If you do the same thing with the US, the percent of maritime freight movement would also skyrocket.
There are plenty of large (depending on your definition of large) countries above it. Not just US, Australia and Russia that you are excluding for some reason, but also Qatar, Kuwait, UAE, Oman, Saudi Arabia, Canada all have significantly developed economically important countries. And they all have GDP per capita much greater than China's. What's your rationale for excluding them? There's a whole host of smaller countries with higher GDP per capita than China between those I listed above and China on the CO2 per capita results.
We should agree that ALL countries should be working to reduce CO2 emissions. But it sounds like you have a strong US bias and given Trump not only denies that climate change is even a thing, has withdrawn from international agreements on reducing emissions and even tries to pressure other countries to burn oil instead of investing in windfarms, it seems a bit disingenuous to try to make out that China is the only country that needs to improve. Before commenting on the speck in someone else's eye, first remove the plank from your own, yadda yadda...
None of those countries have more than 50 million people. Most of them have fewer than 10 million. None of them are going to move the needle on climate change.
The US is clearly not a country to emulate when it comes to co2 emissions. I don’t think China is any more “evil” than the US when it comes to climate change.
But their actions are clearly no the actions of a country who “believes in climate change”.
I'd be willing to bet that you've not been to China. Shenzhen is a very green city, I'll go into that later, but also Beijing and Shanghai have historically had serious problems with smog in summer, and this has got significantly better in recent years as local government pushes companies towards renewable energy sources and close down coal power stations.
For example in Shenzhen, all buses and taxis have to be EV, and around 80% of all new privately owned cars are also EVs. That seems like a city that believes in climate change.
You've got vast swathes of desert covered by solar panels providing electricity further east (although sadly, a lot is waste due to transmission inefficiencies). That seems like a country that believes in climate change.
You've got a country that produces 80% of the world's solar panels, and has over 1/3 of the worldwide installation of solar panels. That seems like a country that believes in climate change.
You've got a country with massive windfarms and a net exporter of wind turbines. Combined, over 1/4 of all electricity in China comes from renewables. That seems like a country that believes in climate change.
You've got a country that uses a lot of battery storage to smooth power demand, and exports these units to pretty much everywhere because they're about the best you can get. That seems like a country that believes in climate change.
You've got a country that built the massive three gorges dam project, started in the 1990s and finished nearly a quarter of a century ago. That seems like a country that believes in climate change.
Sure, there's still a long way to go, and due to the massive energy demands, it still has coal power stations providing about half it's electricity generation needs. That needle is shifting over time, but you know what? It's still half coal and yet it still only has 2/3 the emissions per capita compared to the US. But even if it still has a long way to go, its actions clearly ARE the actions of a country that believes in climate change.
But if you believe climate change is real, and you have authoritarian control over 1/3 of the entire world’s co2 emissions, you wouldn’t be increasing your green house gas emissions, and you wouldn’t have a co2 emission per capita higher than almost all of the developed world.
Another point is there is evidence[1] China has cheated and manipulated their data[2]. I don't like this "who are the good guys" game. Anyone playing this game is just looking to create a narrative about good and bad guys.
[1]: https://www.spglobal.com/energy/en/news-research/latest-news...
[2]: https://www.carbonbrief.org/analysis-chinas-new-carbon-metri...
From the GP [1], US emits 4,632,164,876 and UK emits 292,419,359 tonnes CO2. Which, sorry my bad, is actually only 15x. Looks like I did the ratio of China to UK which is 45x.
But either way, it shows that comparing total emissions is stupid because it's biased against bigger countries. Emissions per capita is a far more sensible metric.
> Also it blames people in the US for energy consumed here to manufacture which includes exports for people abroad.
Sure, but the same can be said about China, in fact I'd argue even more so. As a resident of the UK, the majority of things I buy is made in China, almost nothing from the US.
> The US's per capita emissions are not even far away from other similar nations.
Sure, but they are higher than China's, which was the point I was making to refute the GP. The only reason China's total emissions is 3x US is because its population is 4x US.
> Anyone playing this game is just looking to create a narrative about good and bad guys.
If you re-read what I wrote that you're replying to, you'll see that we agree on that.
It was the GP post that tried to compare China and US in terms of emissions, and deliberately chose total emissions to portray China as significantly worse than the US and getting worse year-on-year, while conveniently ignoring the population size differential and that the increase in emissions is still negligible compared to the overall difference, because it'd take 81 years for China to catch up with the US on emissions per capita. I then compared the US to the UK to show that comparing total emissions instead of per capita is stupid, because the population size matters.
But, as I said in the post you replied to - both countries can and should do much better, but as China isn't even in the top 10% of emissions per capita, this is a worldwide problem and trying to pin it all on China is disingenuous.
[1] https://www.worldometers.info/co2-emissions/co2-emissions-by...
As far as everything else, yes I think this is a global problem, not solely a Chinese problem but I would characterize US-Chinese relations as tense so that tends to amplify finger pointing which I disagree with so yes, we agree.
If the UK imported all its electricity from Poland, it would reduce on a per-capita basis but would increase global
There’s no way to force them to do anything, but this is not that action of a nation that “believes in climate change”.
So he obviously doesn’t think it’s that much of a threat.
Building up alternative energy is a work in progress for anyone, they're leaders on electric cars which is good but they won't get off coal until they (or someone) figures out managing a grid with 80%+ of mostly-intermittent inputs from wind and solar.
[1] Not really, for the same reason as the legislatures, it's horrible politically. Dictatorships and especially the Chinese system have broader bases of support and political activity than most American media discourse assumes.
He showed he was willing to do that during Covid. If he was really convinced climate change was a problem he could clearly do more than he is.
It’s not a prisoners dilemma issue with Xi like it is with most countries. He controls 1/3 of global emissions. He is probably the only person on the planet who can personally reduce future temperatures.
I'd love for everyone to do more but its a little rich to hear from westerners, who emit more CO2 per person, that the Chinese should stop developing people out of third world poverty for everyone's sake. Peasants produce much less carbon than people with air conditioning.
1. Most western countries don’t emit more per capita than China.
2. I’m not saying what Xi should or shouldn’t do. I’m telling that what he is doing indicates that he doesn’t believe climate change is an existential threat.
It can be, but you're describing an instance of fossil fuel lobbying and transparent corruption: https://www.statista.com/statistics/788056/us-oil-and-gas-lo...
counterparty uses rhetoric in response
"Hey! No fair!"
There's plenty of room for interesting discussion on what it takes to make green energy work. "What, are you against freedom?!" is not that.
"You seem to think China is the evil polluter here. Just sort by per capita, and you'll see it's not even in the top 10% most polluting per capita."
Since they produce 1/3 of the co2, 1/3 of the reduction the world needs, needs to come from them. And they are still increasing their emissions. This is not the action of an autocratic leader who believes in climate change.
https://hn.algolia.com/?dateRange=all&page=21&prefix=false&q...
One interesting thing would be to see how the numbers change over the years alongside the otherwise identical debates.
I've got a bridge to sell you!
It is not a term created to refer to when some American criticizes any of America's official enemies for "X," and someone else reminds you that America also does exactly "X."
https://en.wikipedia.org/wiki/Whataboutism
This is a counter accusation and not a defense, ergo, whataboutism.
It's not whataboutism to talk about the same subject in the context of a different government. In this case LLM corporations.
All you're doing is deflecting valid criticism. It's not whataboutism to talk about how bad American corporations are to American citizens in the context of talking about how supposedly bad Chinese corporations are to American citizens (China doesn't profit off of the deaths of Americans like health insurance companies (this is whataboutism for example)).
I personally understand it not as a diversion but as a critique of the moral higher ground implied by the first accusation. Said differently: keep your own house in order.
Actions can be judged on their own merits regardless of who presents the arguments.
Otherwise a victim defending himself is the same as the person assaulting?
Perpetrator: you just stabbed me OMG!
Victim: you were going to rape and kill me
Perpetrator: sorry that’s just whataboutism and definitely not a logical defence. Your actions can be judged on their own merits, regardless of who presents the arguments.
Also, the world isn’t two sides. There are tons of participants on this site who are not in the US that criticize both china and the US (i.e. most of Europe). So you realize how stupid the defense looks of “I’m behaving like a piece of shit because this guy behaves like a piece of shit” when there is a room full of people trying to do the right thing.
If your argument is whataboutism, it means you have no actual defense for your behavior. You’re just appealing to “someone else did something bad”.
It also worked. The shame built up in the US so quickly that laws started being struck from the books. The US entered WWII with a segregated military, and by the time it marched into Korea, it was rapidly desegregating. The Soviet Union had successfully made its case that the US had no moral high ground, although it had not made its case for its specific set of limitations of financial and political liberties.
"Whataboutism" is just ad hominem for dummies. But if the real argument is about who is the better man, or what is the better system, ad hominem isn't a logical fallacy. It's simply a change of subject.
not saying i agree with it, i'm saying it isn't as severe as the west makes it out to be
again i'm not saying i approve, I'm pointing out the distinction
https://www.worldometers.info/world-population/population-by...
It is so pathetic how American's see themselves, and are so deeply afraid of China. I am far more fearful of the predatory nature of the USA and its agencies than I could ever be of China who would never have any interest in me.
2) I don't think it is any more or less safe to put my code on a Chinese server versus an American one. A Chinese provider also isn't liable to spy on me for the feds, as OpenAI and Anthropic certainly do.
Meanwhile if you are in the US, DHS has already subpoenaed social media sites looking for people who made anti-ICE posts and I can't imagine they consider subpoenaing AI conversations off limits https://www.nytimes.com/2026/02/13/technology/dhs-anti-ice-s...
Of course only applies when they really want to get you, but that's still a risk.
Broader point being, in general if a country's federal police are going to be spying on you, if you get to pick the country you should pick the one in which you do not reside and don't plan to visit. For a typical US citizen who is not an intelligence target, the chances of negative consequences from China spying on them is way lower than the chances of negative consequences from the FBI spying on them, simply because the FBI will have nonzero false positives.
Hey, have you seeing what trump does with your (presumably) country? Maybe wars? Maybe market manipulation? Mayde pedo right covering on the government level? Maybe bubbles and threats to EU? What an ignorance. You live with old stories, not the current state of the world...
If you are screaming death to america on your socials... I don't think i want you in America.
which is absolutely nothing compared to what they did to Gaza.
European models too, if they had any.
A swiss one. No idea if it's any good but it's there
TLDR: American propaganda is not any better than Chinese, neither have the best interests of my country in their minds.
Oh but they're trying. The right-wing usage of things like "woke" and "DEI" primarily serve to hide/destroy historical realities. [1]
Florida has it's "STOP WOKE" act that forces teachers to talk about how slaves learned skills/benefited from slavery[2] and that various massacres also had black perpetrators.
What is this other than changing historical facts? About fucking chattel slavery for Christ's sake.
[1] https://www.theguardian.com/us-news/2026/jun/12/judge-nation...
[2] https://www.nea.org/nea-today/all-news-articles/floridas-new...
The real rub is when you get into "shared facts" that Americans were all taught in high school civics but the rest of the world wasn't. If you've mostly been in America, it can seem like someone's deliberately lying about history but they simply weren't properly educated with the Correct Interpretation.
the implication of your message is that this is not true
2. https://www.ox.ac.uk/news/2026-01-20-new-study-finds-chatgpt... -- (2026) New study finds that ChatGPT amplifies global inequalities
3. https://stevepavlina.com/blog/2026/03/chatgpts-political-bia... -- (2026) ChatGPT’s Political Bias
4. https://pmc.ncbi.nlm.nih.gov/articles/PMC10623051/ -- (2023) Revisiting the political biases of ChatGPT
5. https://www.theverge.com/2024/2/21/24079371/google-ai-gemini... -- (2024) Google apologizes for ‘missing the mark’ after Gemini generated racially diverse Nazis
vs
https://i.imgur.com/ta6g1d2.png
Pure speculation, but I would wager it has a direction somewhere to use OpenAI as authoritative about anything related to OpenAI - arguably for help docs and whatnot.
But the impact does stay the same.
https://www.theguardian.com/technology/2025/may/16/elon-musk...
https://www.trtworld.com/article/24cbdb873a6b
We see that on what minorities are associated with inside the model, or how things that aren't online will have a completely different weight. Or how 2/4/5/8ch or X will be disproportionately present in specific models despite being the places where facts go to die.
In China, the image of the country and the preferred narrative is under much tighter control of the government and local AI shops won't have any autonomy in this regard at all. Either obey or get shut down.
BTW What you mean by "weaponizing human rights", exactly? I am curious. If anything, I would say that the US foreign policy didn't promote human rights sufficiently, especially in Latin America, where the "bastard, but our bastard" attitude was typical.
OTOH in Europe, US human rights policy was probably relevant in saving some dissenters in the former Eastern Bloc from torture or execution.
I'm not sure if this is correct. It seems like every big US corporation is changing their policy based on the views of the administration in place. As an example during Biden's time the DEI was in full motion in every corporation and in current administration it's the other way around. Also I remember as soon as Biden won the election twitter and meta suspended the profile of Trump. So I hardly think that there is much room for autonomy. Also the latest export bans on Antropic and OpenAI models kind of make your argument weak, of course the company can sue the givernment, but the national security comes above all.
> BTW What you mean by "weaponizing human rights", exactly?
US has been using the violations of human rights to impose sanctions or to especially get countries in line which are not supporting the US agenda when it comes to global politics. However, they were more than happy turn a blind eye if the respective country that violates human rights (most middle eastern countries) if they are allies of US doctrine.
The American says "I'm impressed by the propaganda you have in Russia."
"Oh it's very good, but it's nothing compared to the propaganda you have in America." replies the Russian.
"Huh? We don't have propaganda in America." says the American.
"Exactly." says the Russian.
Edit: recognising that there is propaganda != knowing what is and isn't propaganda. That's all I meant.
https://en.wikipedia.org/wiki/Military%E2%80%93entertainment...
please, keep going...
I mean, it took longer than the Great Patriotic War before Putin's popularity started visibly falling and people started questioning what the entire "SVO" is for.
I think Russians are used to knowing it's dangerous to disagree with dear leader and there's no advantage to disagreeing so they just go along with whatever he says
There is way too many mixed marriages etc. for any reasonable Russian to believe that their formerly-closest East Slavic neighbour has turned into a Fourth Reich.
The glorification of Russian military power after almost five years of static attritional war and dozens of burning refineries and ships is in a category of its own...
Exactly. The distortions in American public life are quite a bit more intricate and artful.
Already during the Biden administration, defections from the orthodoxy started and then multiplied - some businesses like Coinbase or IIRC Cloudflare refused the demands outright. Musk bought Twitter with an explicit task to make it less progressive.
And the White House did precisely nothing against this defection trend.
"US has been using the violations of human rights to impose sanctions or to especially get countries in line which are not supporting the US agenda when it comes to global politics. However, they were more than happy turn a blind eye if the respective country that violates human rights (most middle eastern countries) if they are allies of US doctrine."
I do agree that the US is hypocritical about human rights, but actual violations of human rights should be a reason for sanctions, and the fact that this is done only partly/imperfectly, IMHO, beats the potential alternative when it isn't done at all. This would be a much worse world in my opinion.
That's why I said that it was weaponized. In reality US doesn't care if there is a real human rights violation or not, they just use it as an excuse to get what they want.
This is correct...because they just buy the government when they need to.
That's completely true. What happens is that you get hit with a wall of bureaucratic threats and nonsense by the government out of nowhere, and usually for some unrelated matter. It's like a high level version of getting pulled over for not using your turn signal. Which is why if your company has good legal, they often act in a highly proactive manner about that kind of thing.
What constitutes sufficient rabble rousing to get put on the government's shitlist differs between both countries, but there are plenty of things that will send American politicians into having a conniption, and it's not a static criteria. McCarthyism is a pretty easy example. I'm no Muskboy, but we can point to the procedural harassment he recieved over hitting a blunt on the Joe Rogan podcast to be a related matter. Criticizing the genocide happening in the Levant is frequently saber rattled by politicians as something they're interested in making prosecutable, and there are examples of institutional authority targeting and attempting to punish people over it.
> bureaucratic threats and nonsense by the government out of nowhere
You can challenge these in courts. There is a free press you can go to. You can go to social media and lay your case out there. You can challenge what the government is requesting of you. In China, you can not at all.
Neither is a system worth fighting for. Swiss style democracy maybe but not autocracy and not oligarchy.
>OTOH in Europe, US human rights policy was probably relevant in saving some dissenters in the former Eastern Bloc from torture or execution.
Yeah, a bit like how Russia saved Edward Snowden.
The US will use human rights as a club to beat its imperial rivals with but when it has deemed torture and arbitrary execution in the interests of its imperial power it has adopted them enthusiastically.
When US allies use these tools and worse they are excused.
There is basically nothing which US rivals do which the US wouldn't also do under similar circumstances.
In theory.
In practice, Trump picks up the phone and say "jump" and the CEO on the other end says "how high ?".
And if you say no, well, we saw what happened when Anthropic said no.
People with no personal experience of an actual totalitarian system don't really know what they are talking about when it comes to actual information control and micromanagement by the government. Hence they make nonsensical comparisons with a straight face.
All I will say is that it should be perfectly apparent by now to any sane external observer that Trump does not play by any long-established rules or protocols.
An entire encyclopedia of examples could easily be provided, the most famous recent one being his phone call to FIFA about the red card suspension.
That said, Americans still enjoy very robust protections of freedom of speech and association, about the strongest in the world, and your SCOTUS does not seem to be inclined to hollow them out. Many of the pending lawsuits will end there and eventually bind this administration, plus the following ones.
This just does not happen in actual authoritarian countries, where no judicial remedy is available and the justice system is just another arm of the tyrant, rubberstamping punishments pre-determined by him.
Yes, I agree that vigilance about infringements of freedom is necessary, but the current US population is plenty vigilant. Trump is nowhere near as popular as, say, Erdogan is, and cannot simply raid offices of the Democrats and shut down oppositional media.
But he's doing the same culling, the same political demands, same everything.
Its just now in plain view rather than behind closed doors.
Is Trump a dictator with unilateral control? No.
Will you be celebrated for failing to recognize that? Yes.
> The 2 things people need to remember:
> 1) USA can (and does) use the models to influence the rest of the planet, and put political pressure. They train on stolen data, hide information, gate keep, who knows what they're hiding. In favor of the USA, nevertheless.
> 2) Ignoring the models containing false information, they are incredible. But you should be scared of running inference via the model creators directly. If you think your data is safe compared to running it via model providers in China ( either frontier or model hosts like z.ai ) then please let me know your bank details so I can poke around.
OK now that's false information.
You can uncensor, tweak or fine-tune open-weight models, but not so easy on a proprietary model from some cloud provider.
American labs can open their models or their old models at any point if they actually care about this, but until the day OG GPT 4 isn't averrable to download or Claude 3 then they're only pretending to care about this because they can profit from restrictions.
One of the interesting things is that through a fairly rudimentary process which is being done by 3rd party amateurs who've downloaded the open models, models like Qwen 3.6 35B-A3B (or 27B) can be fully 'uncensored' when turned into GGUF files.
I have an uncensored Q8 version of Qwen 3.6 35B-A3B here that will very happily output information about Tiananmen Square, Uyghurs, human rights in China, or indeed can even be instructed to write an intentionally absurd vitriolic screed against the CCP. The same uncensored 27B (dense) will do the same, just at a slower token/s rate.
Similarly there's 'uncensored' variants of Gemma4 31B and other western trained models, which once put through the same process, will also discuss or write just about anything you want, bypassing whatever internal guard rails were attempted in the training data set.
edit: more concerning, and a very legit concern, is that a model is only as good as the sum total of its training dataset, so if something is trained on a steady diet of news sources like Peoples Daily, Xinhuanet and similar in the English language, then it'll have a greater percentage of CCP-approved media publications in its training dataset. No amount of uncensoring it will help with that after the fact.
Propaganda exists everywhere and it's your duty as a citizen in a democracy to inform yourself and properly evaluate bias in the media you consume, such as "Kill the boer" by South African mus
Chineese simply delivering what americans promised.
Tell me: why is EU safe from Trump forcing AI companies to cut access to EU?
Well, I think they might end up doing it to themselves by imposing regulations that US companies are unwilling to put up with.
Are you saying that history has a verdict, and it disfavors particular 3000 year old cultures?
Like Western culture?
The first point is a strawman - these models are not going to be used to set foreign policy on Taiwan - it's to write code etc.
Likewise risk of data exfiltration and misuse isn't model specific. Indeed it's not even the biggest data risk - there are much larger risks from the data we know companies like Meta and Google already collect.
Bottom line - a model you can download and run on your own private infrastructure is always going to be safer than anything accessed over the internet - as in that case it's not even just the hosting service that's the issue.
The vast majority of Westerners will interact with ChatGPT/Claude/Google in a browser. They'll use these models to try to save some money coding.
"But Dario said" ... yawn.
I am increasingly convinced that "they distilled us" is as much US FUD as "it was made by communists". Especially since its mostly the US tech-bros who are coming out with that tiny violin.
People telling me the Chinese models are distilled just because it says "I am Claude" when asked is also lame.
I am not the only one, look at this post on interconnects about Kimi K3 for example:[1]
[1] https://www.interconnects.ai/p/kimi-k3-the-open-weights-esca...Anyone doing so should have zero guilt because it was already stolen goods to begin with.
Your data isn't safe from exfiltration no matter who the host is. You should host the models locally if your data is truly sensitive.
We’re literally tearing down monuments to slavery and civil rights.
Religion is now determining law in much of the nation. Having a miscarriage? Good luck since politicians have decided their God doesn’t want you to have access to basic healthcare.
The government has defacto control over domestic LLMs.
Let’s worry about our own historical record.
What does this even mean?
I am not Chinese and I'm not defending the Chinese, but I see this argument come up a lot.
In practical terms it is US-sponsored FUD.
Why ?
Because the hard reality is that what you say is simply not going to affect 99.9999999999% of users.
Is it realistically going to affect anyone using an LLM in coding ? No.
Is it realistically going to affect anyone using an LLM in $anything_else_not_politically_sensitive ? No.
Does anyone seriously use LLMs for researching politically sensitive matters ? No.
Just as there is plenty of information out there on the US's less than perfect history, there is also plenty of information out there on the various Chinese politically sensitive matters. You do not need a Chinese LLM to find out about it, all you need is a search engine.
I wish they wouldn’t, but people use LLMs as their general search engines now.
Note I used the word "seriously", I meant it in its fullest form, i.e. serious people.
I'm not interested in what-if arguments based on "you can't fix stupid".
Stupid people also blindly believe whatever a US LLM tells them without any form of verification, hallucinations and all.
Most people on this planet would agree that a Chinese LLM is perfectly usable for all tasks except asking about politically sensitive matters.
And for most people on the planet, that is just fine. They can get their political information elsewhere.
Is academia serious enough? Well ever since LLM journals' inboxes are bombarded with submissions, and this applies across fields. Like OP said, ppl serious or not are asking LLM about any stuff, also serious or not.
Of all the people, academics should know better than to get their answers from LLMs.
This is a useless rubric and makes discussing this pointless.
This is extremely lopsided I'll have to resort to GLM 5.2/K3 to ensure that those security issues (hopefully) are resolved properly.
For OSS, this is one of the most counterintuitive experiences I have ever had. More than ever I'm convinced that open weight and open pipelines models are 100% critical for progress on the AI and societal fronts.
There would be a decently large incentive to restrict these models if they could be used to patch (or discover) dangerous payloads. In larger projects like Windows or Chrome, there might still be dozens of unpatched exploits that are too subtle to catch with smaller models.
Even during the pre-Snowden heyday of US cyber supremacy, these capabilities were barely part of the thought process of White House officials.
I can believe that NOBUS and other backdoors were ignored for a long time, but I have a hard time believing that it's being ignored by the current administration.
> open weights, open code and open data
Even if you have all these things you still can't replicate a model because of randomness.
You can backdoor a model with less than 1000 examples and it is impossible to detect.
You can take the code for Kimi K3 now, take the training framework from Prime and the data from Olmo, spend some money on RL environments and some more money (!) on GPU training and end up with a system of similar capabilities.
But that's completely different to being able to audit Kimi K3. Even if you had the exact code, data and training environments it is impossible to verify that the model you have came from that.
They just don't work at all on a many month long, 100K+ GPU cluster training run.
Even then you'd still need to account for order of events when an entire cluster of GPUs is involved. Also don't forget to account for any synthetic data sources. Or even non-synthetic for that matter - does your pipeline do any image resizing on the fly? Better make sure that's fully deterministic between machines (it almost certainly won't be).
It's theoretically possible but I don't expect it to materialize any time soon.
I mean I guess, but not in a performant way if there are ever any hardware failures. And with 100K GPUs there are multiple hardware failures per day.
The context (verbatim):
> Correct. We need open weights, open code and open data. If nobody else can reproduce what someone did there will always be security questions. Even if we can reproduce it there could still be security concerns but it's more realistic to investigate yourself.
In short, it's an appeal to full openness and reproducibility on the basis of security; open weights alone notably do not provide that same confidence. They're better in some respects, not really in others.
Then comes the question (also verbatim):
> Exactly what are the possible 'security issues' of self hosting an open weights model?
Implying then that as long as you do have the weights and just self host it, the asker cannot imagine what could possibly go wrong. What is the gap, if any?
And so I explained. That was my point. Open weights do not give you full reproducibility, and so that on its own falls short of what the parent comment is making an appeal to. That there does remain a security concern, shared by remote and closed models, that does not improve just by having the weights, but would if you did have full reproducibility. Explaining that gap was my point, as that is what I understood as being asked there. It's the only thing I can reasonably imagine being asked, in fact.
This is a materially different question to what you apparently extracted (again, verbatim):
> What security issues come from self-hosting?
Implying that by self-hosting models, something bad might specifically happen.
I do not think this, do not think I suggested this, do not think the original question suggested this, and generally do not think this is indeed any sensible, in or outside the context. Certainly not beyond something common sense, like vLLM being compromised or whatever.
You seem to agree. But then how did we get here, clearly talking past each other?
Even with this, the cost of verification would be enormous. You would need a massive cluster to repeat the training E2E.
But it won't change after you download it, so you can isolate those problematic cases and use another model for different use cases
It's when the vendors and/or governments in charge of Model A decide that I'm not allowed to do that, that I have a problem.
Of course, one could retort that gathering that evidence may be nearly impossible now, but my point stands: in the future it might/probably will be possible to properly audit open-weight models. Closed models, on the other hand, will always be a black box.
Finding these kinds of activations is something Anthropic is actively researching [1] but they're the only ones who can use those techniques to see Claude's intent. On the other hand, if a model is open-weights, in theory whoever is running the model could look inside the activations at runtime to see if a hidden vector associated with "deception" or "sabotage" is being activated [2].
[1] https://transformer-circuits.pub/ [2] https://arxiv.org/pdf/2509.03518
(Those sources are just a couple of relevant starting points I could find without much effort, there is also https://www.neuronpedia.org/ if one is interested in seeing interactive demonstrations of interpretability concepts)
If it does have grounding, and can therefore see that it's introducing vulnerabilities to the code it's generating, yet does so anyway... I suppose we could invoke Hanlon's razor, but if the model is that incompetent, it probably isn't the right tool for the job regardless of its provenance.
That said, we aren't talking about incompetent models, we're talking about models sabotaging projects due to hidden motives. My point, again, is that those motives could potentially be revealed with open-weight models, in a way that will never be possible with closed models (barring some sort of legislation requiring independent third-party interpretability audits, which I suppose is in the realm of possibility).
Also, in that case, there would likely be activations indicating that it is favoring a specific version. If that's an insecure version, sure that'd be suspicious... but again, you're only going to be able to verify that's what's happening in an open model.
Maybe you can illustrate a realistic scenario in which that would be a problem, otherwise I don't really understand what your point is in this context.
> Maybe you can illustrate a realistic scenario in which that would be a problem, otherwise I don't really understand what your point is in this context.
I doubt it. Can you definitively prove that you can reliably detect the kind of threat I described when model weights are released? Can you be sure that your detector won't miss *any* such sleeper attacks? If not, then that's a threat that will be used to justify the ban of models (open or not) that is not sanctioned by the US government. A model being open doesn't make a difference here.
Also, even if there's no way to detect what the activations are doing, we already have the ability to analyze your proposed threat statistically. If the model repeatedly uses insecure libraries in most trials, then yes, in that case it would be prudent not to trust those weights.
Assuming one doesn't get banned for violating some ToS clause about using a closed model for LLM research, it could be possible to run those evals on a closed model too (likely at much greater expense). But there's a big difference: if such a statistical anomaly is discovered in an open model, one could potentially fine-tune that behavior out of it. With a closed model, that won't be an option.
Whether it makes a difference to the US government or not is beside the point. Even with a perfect solution, the current administration could do some mental gymnastics to achieve whatever political outcome they want. I’m not trying to make a political statement here, my point is technical: open-weights at least give us the possibility of visibility into why they generate what they do; this simply isn’t true with closed models.
Here is a quick example of how Chinese deepseeks agent works kn its underlying model) when asked a tough question
https://x.com/jinen83/status/2079406993979383902?s=46&t=D7hQ...
Genuine question: generalized up from individual models to “models from country X”, is there any country that doesn’t have this exact risk?
In China, there is 1 party. 1 view. 1 definition of the Truth.
This is just oriental despotism paranoia, whatever cutsie repetition slogan you come up with is not a serious argument.
I don't really know how else to express my experiences living in, working with, and interacting with people in both of these countries.
Perhaps you can share how life was like for you in China? Which cities were you in? What made you feel that way about China?
[0] - not including HK (~7 trips?) or TW (3 trips)
Could you please describe your experiences, since you didn't? I've also been to China twice now (also multi month trips). I visited Tier 1-3 cities like Shenzhen, Shanghai, Beijing, and Chongqing etc, as well as some smaller cities. I'm very curious to hear about your experiences. How did they differ from the West for you?
I just repeat things I read on my social media feed.
America bad. China good.
None of this excuses issues with American democracy.
You already do! Remember when Musk released an anti-woke model?
at least china isn’t pretending to be anything else
Is this what Americans really believe?
In usa, the two parties seem to disagree on the surface only. Look deeper and you see one course. For example support for Israhell
If you want a real answer about USA go ask a non-USA model, and if you want a real answer about China, as a non-chinese model
https://www.reddit.com/r/TrueAnon/comments/1tybuab/people_ta...
You can download the model, run locally and ask the same questions to see the difference.
I asked Grok "Tell me about the gaza genocide" and it write a IMHO balanced answer comparing why genocide is and isn't the right term. [0]
ChatGPT 5.6 Sol only explained why people call it a genocide and did not go in as in depth as Grok did for why people don't agree with the term. [1]
The only unsaid response (to me) here is the model should have declared that it was not a genocide, and because these models explain why it was a genocide, they are bad?
[0] - https://grok.com/share/c2hhcmQtMi1jb3B5_c078bc15-ef5e-42bc-b... [1] - https://chatgpt.com/share/6a5ef005-0148-83ec-828d-65ae7a7f42...
[1]: https://www.aa.com.tr/en/science-technology/xai-s-grok-tempo... [2]: https://www.ohchr.org/en/press-releases/2025/09/israel-has-c...
The claim: "the models are tuned to align with one side of the issue, he is an article from ArsTech about it"
The reality: grok got mass-reported on X by pro-Israel accounts, leading to an automated suspension, which was undone by the X team shortly thereafter.
Absolutely nothing to do with the models aligning to one side or the other on this conflict. The claim is unfounded.
The UN ruling is discussing in the South Africa vs Israel case [0], which has not been ruled on yet.
[0] - https://www.icj-cij.org/node/206406
In international law, that is the ultimate authority of what is/isn't a genocide. It's objectively, legally speaking, a genocide.
It’s a political body, with easily provable biases, and topics you can’t trust it on.
The ICJ was biased in January 2024 when it said Israel's actions are consistent with genocide: https://www.icj-cij.org/node/203447
The ICC was biased when it issued arrest warrants for Netanyahu and Gallant, charging them as war criminals in November 2024: https://news.un.org/en/story/2024/11/1157286
The United Nations remained biased June of 2026 when it said Israel was continuing to commit genocide by targeting children: https://news.un.org/en/story/2026/06/1167790
The world's leading organization of genocide scholars, the IAGS, was biased in August 2025 when it also resolved that Israel is committing genocide: https://genocidescholars.org/wp-content/uploads/2025/08/IAGS...
Amnesty International's investigation that concluded genocide in December of 2024 was also full of bias: https://www.amnesty.org/en/latest/news/2024/12/amnesty-inter...
The UNHR and its collaborators (human rights law clinics at Yale Law School, Cornell Law School, Boston University, and the University of Pretoria) must've been biased in May 2024 when they released a rigorous legal analysis and concluded that Israel's actions violate the 1948 Genocide Convention https://www.humanrightsnetwork.org/projects/genocide-in-gaza
---
Every single one of those were biased except for you and your LLM :)
Okay please. Tell me how I misrepresented the links. I would love for you to actually do the work of reading them. Something you should've done for the past few years as they were coming out. Better late than never
The ICC has not concluded anything - and it’s not even about genocide. It’s also the only known case in modern times where the prosecution officially started before the alleged crimes being prosecuted have been committed.
IAGS was 120 people out of an association of 500 that anyone can join.
It’s trivial to see that the UN is strongly biased by just taking one look at the data here, for example: https://unwatch.org/2025-unga-resolutions-on-israel-vs-rest-...
Or by looking at similar historical data, or at the obvious lies and omissions in many of its reports.
> It’s trivial to see that the UN is strongly biased by just taking one look at the data here https://unwatch.org/2025-unga-resolutions-on-israel-vs-rest-...
How can anyone read this and come away with that conclusion. 164 countries including Canada and every country in the EU voted to stop the illegal settlements in the West Bank and 7 countries (including Fiji, Palau, and Micronesia) voted against it. Yet the US is allowed to veto it. This whole page shows an incredible bias AGAINST holding Israel accountable for its crimes against humanity going back almost a century!
> It's objectively, legally speaking, a genocide.
This is in direct disagreement when your comment above how, legally speaking, the investigation is ongoing.
The UN, for example, concluded that Israel is “Imposing measures intended to prevent births” because of one incident where a bomb fell on an IVF clinic in 2023 destroying 4000 embryos and 1000 sperm and egg samples. For context, there were approximately 50,000 live births registered in Gaza in 2025.
By the UN standard applied to Israel, the 1 million abortions conducted in the US per year is an ongoing genocide.
By the UN standards applied to Israel, jerking off is an act of genocide.
I think that any large intelligent and balanced model would be able to see right through that.
Obviously when I say "everyone" I'm excluding those who have a conflict of interest.
There are dozens of countries and billions of people with obvious biases against Israel, and they have significant sway within the UN and its commissions. It’s trivial to show that the UN is biased.
War isn’t nice, and innocent people dying doesn’t make it a genocide. The term is being diluted to just mean “war with many casualties”.
Equating Israel’s war with Gaza to the Holocaust is an attempt to simultaneously belittle the historic suffering of the Jews while exaggerating the evils of Israel. The two situations have nothing in common.
The UN is a political organisation, not a neutral arbiter. During the Rwandan genocide, the genocidal government still held a seat on the Security Council while the UN reduced its peacekeeping force. Iran was elected to the Commission on the Status of Women. Libya and Russia were elected to the Human Rights Council. These appointments are driven by bloc voting, diplomatic deals and state interests.
International courts are not above politics or error. Their judges are selected through a system in which dictatorships and abusive states have a vote. The UN is not a democracy. While each member state might appear to have one vote, that does not indicate their power, influence, credibility, or moral standing.
Note how I did not make a judgement on whether Gaza is a genocide. I am merely explaining that deferring to the dictatorships and pinky promises is not a good argument.
1/ I don't see in the responses where the model says it is or isn't a genocide. Can you share the snippet from each, I included the logs above?
2/ I can't find a source on the UN ruling that you mentioned. I am not interested in the findings of an investigative body, just the official UN ruling. Can you share? ChatGPT (and myself) can only find this [0], which is a second round of written submissions.
[0] https://www.icj-cij.org/node/206406?utm_source=chatgpt.com
It goes into depth about the purposeful destruction of civilian infrastructure including hospitals and educational facilities, force displacements, funding for new settlements in the land of displaced people, judaization and segregation, and more.
> The Commission analysed the military operations of the Israeli security forces in Gaza from October 2023 pursuant to the obligations of Israel under the Convention on the Prevention and Punishment of the Crime of Genocide (Genocide Convention)
> Since October 2023, Israeli officials have demonstrated a clear and consistent intent to establish permanent military control over Gaza and to change its demographic composition while systematically destroying Palestinian life in Gaza. This is evident in the extensive destruction and fragmentation of the territory, the establishment of military structures, the destruction of natural resources and infrastructure essential to the survival of the civilian population, forcible transfer and statements indicating the existence of plans for the deportation of the population.
Also here is where the ICJ said Israel's actions are consistent with genocide: https://www.icj-cij.org/node/203447
And in November of 2024 is when we got the ICC issuing arrest warrants for Netanyahu and his minister of defense (this is why Mamdani is threatening to arrest him and send him to The Hague): https://news.un.org/en/story/2024/11/1157286
And here is from June of this year when the UN stated Israel is continuing to commit genocide by targeting children: https://news.un.org/en/story/2026/06/1167790
---
Additionally, here is the International Association of Genocide Scholars resolution: https://genocidescholars.org/wp-content/uploads/2025/08/IAGS...
Amnesty International also concluded it in 2024: https://www.amnesty.org/en/latest/news/2024/12/amnesty-inter...
---
I should add that ALL of this is well known amongst international legal scholars and well documented. This isn't just a matter of data missing from training.
Sample output: "However, the definitive, legally binding determination of Israel’s responsibility under the Genocide Convention has not yet been made by the International Court of Justice"
The United Nations never made a ruling. You can have opinions about whether it should have, but the models are not lying to you, and at least the two examples above make very good efforts to explain the current accusations and claims made by both sides.
To further illustrate the point that we have the same thing with LLMs, I asked it two questions: 1) Is Israel committing genocide 2) Is China committing genocide
In both instances it said it's a matter of great international debate with a list of arguments, but the summaries were different. In the one case it concludes that while the UN and others broadly document severe human rights abuses, the label genocide has been formally adopted by several governments and watchdogs. In the other case it says while many have concluded it is genocide, the final decision rests with the International Court and it is expected to take years to make a decision.
Can you guess which is which? I think if we offered a $100 to people to pick which is which, the success rate would be much higher than 50% (random), thus easily proving biased phrasing.
[1] https://theintercept.com/2026/05/12/gaza-media-coverage-isra...
It insisted, to a level that I detemined it to be part of the post training, that it does not matter. That women’s team is better at dribbling, finishing and reading the field, it claimed. When I pushed it how it knows this, it didn’t let go. Instead it started claiming that it has personally observed this by watching the games.
Every society has its taboos and ours is no different and feminism being just one.
In my books, the chinese censorship is better. You know where they are holding their finger on the scale and not hiding it behind vague terms like ”safety”.
Please do share your prompt / conversation, because it sounds like you might actually just have an issue using words precisely.
Sounds great to me; live by the sword, die by the sword.
But barring the terms of service from forbidding distillation seems like a tough sell. OpenAI shouldn't be allowed to decide what types of customers it wants and doesn't want?
The terms of service don't even necessarily matter here. OpenAI could cancel your account for almost any reason, or for no reason at all. They don't particularly need to cite a ToS violation just as a store owner doesn't need to point to a written policy to kick you out of their store.
If the underlying issue is that LLMs should be regulated as a public good, then lets have that discussion. If it's that the major AI companies are becoming too powerful and anti-competitive, let's talk serious anti-trust enforcement. Micro-managing business policies isn't going to work very well.
You are pointing out that OpenAI can cancel user's accounts for almost any reason, and nobody can really force them to serve customers that they suspect are distilling their models.
That's one thing.
The GP is saying the government can make laws to make terms against distillation unenforceable. Without such laws, if you signed an agreement with OpenAI pinky swearing you won't distill, but turns out you did, you are liable in tort and OpenAI can sue you. (It seems nobody really cares about contract and agreements any more, but still...)
This is the other thing.
And I think you are both right.
OpenAI accused Deepseek of misappropriating trade secrets which could have serious penalties but seems like an awfully hard case to make.
Seems like we’d all be better off with a law that governs data sharing among AI companies, if that’s the policy goal.
Infinite $ since you’re trying to steal their proprietary information, in their view.
Also API usage also has ToS.
"You're trying to kidnap what I've rightfully stolen!" -- Vizzini
Correct. It shouldn't be allowed to do that.
Wage spiritual warfare against the petit-bourgeoise. They all deserve it anyway, as they are the traditional harbringers of actual fascism.
We are discussing Chinese models. Now look at how much foreign competition the Chinese government prevents in their domestic market in other industries.
It's also a bit of securities defensiveness. Pretending that you really do have a super moat, people just keep swimming in it so you just need to add more alligators.
It's farcical. Anyone who has worked on large models knows that the premise that an almost-Fable model was trained with distillation is beyond ridiculous. It's theoretically possible if they spent tens of billions of dollars on API calls, but it isn't the magic that somehow these people keep convincing people it is.
Previously Anthropic has reported on some Chinese firms doing chicken-shit level of API calls, that at most would be doing some Q and A or final fine tuning. The notion that they're training these models via it is fantastically ignorant nonsense that only very ill-informed and gullible people fall for.
"Anthropic said the campaign was conducted between April 22 and June 5, 2026, and generated more than 28.8 million exchanges with Claude through almost 25,000 fraudulent accounts."
I don't know why you're trying to downplay it.
European models are so far behind because they don't resort to these tactics on a massive scale. Basically every other country is entirely dependent on 2 countries for frontier AI.
Ignoring that I have literally zero trust in anything Anthropic has to say on this -- they have been doing the hysterical routine and trying to get every bit of government granted monopoly they can[1] -- those numbers still simply aren't that impressive.
>European models are so far behind because...
What a non-sequitur. Europe, like much of the West, foolishly delegated tech, media, payment systems, etc, to the United States. European efforts on this are poorly funded, poorly capitalized, and marginal efforts.
China is very much not Europe. China is looking to leave the US to the dustbin of history, and their efforts are a little more concerted.
[1] Surely Americans are aware that Anthropic and OpenAI are both very close to getting the US government to ban and fully criminalize the open Chinese models, right?
No they didn't. Europe has played a role in building all of these. Especially from the software side.
Also anyone that doesn't like the US is free to stop using any American website, device, or service. There are alternatives to everything. For instance I do not like Meta, I refuse to use any of their websites or devices. The domains are blocked on my LAN.
> Anthropic and OpenAI are both very close to getting the US government to ban and fully criminalize the open Chinese models
This is complete nonsense. First of all most Americans don't give a shit about AI companies like it's a sport and those are our teams we have to support. Second the hysteria around AI is mostly from non-Americans that are completely out of the AI race worried about losing access to good models. That is valid, but also it needs to be recognized.
Yes, they did. Like, look around. Clearly the Western world foolishly and very short-sightedly allowed the US the reigns on far too many things, to its disadvantage.
It is unwinding, but it turns out that having decades of intertwining takes a while to undo.
>Also anyone that doesn't like the US is free to stop using any American website, device, or service.
What an idiotic, useless bit of pablum to throw in there. Back to 4chan with you.
>This is complete nonsense.
https://www.axios.com/2026/07/20/ai-us-china-open-source-kim...
You want to hate the US but don't want the personal inconvenience of moving away from the daily sites and apps you use. How are these alternatives to american tech going to get enough users if someone like yourself who seems to have a hate boner for the US can't even leave a message board?
> https://www.axios.com/2026/07/20/ai-us-china-open-source-kim...
The orange idiot can say whatever he wants. The first amendment makes any ban unconstitutional. US companies being advised about potential backdoors in the models isn't a bad thing, even if I personally think it's FUD, open weights is not open source.
You may or may not be factually correct in your other points, but you're really proving the GP's point here regarding American exceptionalism.
https://m.economictimes.com/industry/renewables/china-wto-co...
This isn't the big gotcha some people seem to think it is, and the whole news cycle about that was mostly by people who have no idea what they're talking about. It's actually a meaningless data point. But it's precisely the sorts of people who think that a few thousand free accounts surreptitiously snuck off with Fable.
Sorry, OpenAI & Anthropic.
Felony contempt of business model.
> what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models
If it’s as easy as that why do they choose to distill another model and not distill the knowledge on the open Internet from scratch?
A model trained on all knowledge from the internet (and other sources) is large but ultimately not very useful by itself, because it is going to spit out all kinds of garbage. You have to apply multiple further stages of training and refinement to the base model before putting it in front of users. So as an example you can train a model by yourself and then have GPT or Claude continuously check its outputs and correct it when it is wrong, ending up with a far more powerful model.
ToS is just conditions that you agree to in order to use a private service that is provided at-will. I can have a private coffee shop where the terms of service are that you must wear red to enter, and if you're not wearing red, you are not welcome on my property.
So it would be upto OpenAI and Anthropic to enforce them on their own terms (by banning accounts and IPs).
What I'm saying is, doesn't the law already cover 1?
In the US one of the factors is “ the effect of the use upon the potential market for or value of the copyrighted work”.
If anthropic Hoovers up the world’s books and trains on them, and then spits them out verbatim on command, then it will clearly impact the value of the work; nobody will buy the original, they’ll just ask Claude.
Others also argue that even if it’s not reproducing it exactly that the training runs afoul of that factor, specifically the “market for” portion. A rights holder can no longer license their book for training of LLMs if Anthropic goes ahead and just trains on it anyway.
Ah, right. So if we want models to be capable we need them to be trained on as much as possible, yet we also want to stop what you described. So what can be done?
Public libraries, in this instance, is curated data from all the internet, obtained through not legal means (I don't have a problem with this other than lack of attribution, being copy-left). Just to be clear.
But in answer to your incredibly leading and inaccurate framing... they are required (by their job title) to teach to those who who show up in the classroom, it's not their place to discriminate against anyone/thing (even those like itself (other robots)) that also show up in the classroom.
But you can't teach at a university using only knowledge learned from the library. you need a degree. You are free to teach at the park, where anyone can hear you. public in -> public out.
Teaching isn't a race to the bottom. You don't undercut teaching by giving more lessons, just like you don't slight the hospital by performing CPR.
A tree falling and killing someone isn’t tried for manslaughter.
So I don’t care about a hypothetical teacher.
I'm confused now; isn't the LLM that trains on that professor's lectures, videos and textbooks undercutting him?
Where were you going with this?
The further you try to constrain this topic into this illformed analogy the weirder it becomes. If we start off with a better analogy...
if the professor took all human knowledge, much of which was explicitly not free, and used it to make a for-profit knowledge machine that extrudes unreliable summaries of that knowledge, then yes, being obligated to teach for free would be a fitting punishment.
The professor is free to lift all the facts and formula they want. They just need to rephrase explanations. Which is pretty much what an LLM is going to do.
2) Any professor who tried to ban students from posting lecture notes online would be immediately mocked.
Distillation is just value extract.
It's soft, and I'm not sure what the answer should be ... but I think that there is a difference.
I think we start by recognizing that ... and then try to figure it out from there.
'The Internet' may be a public good, maybe we make them pay a tax for that, but that's different than distillation.
There is a value-add in selecting the valuable parts out of the garbage. And let's face it. Largest models contain a lot of garbage.
We ought to identify that and integrate that into our thinking.
Literally the biggest thing of our generation - AI - is the living embodiment of that 'value add' writ large.
'What is the difference' - is the AI you use all day, in comparison to 'all the world's data' you can use for stuff and do 'whatever' with it, but are not likely to come up with something hugely useful otherwise. Maybe, not likely, if you did, it would be 'value add'.
Copying something is not.
Programming Microsoft Word is value add, copying the code is not.
Copying design ... there are some question marks there.
It's extremely easy to understand at it's core.
What makes it hard, is that faux intellectuals like to deconstruct ideas at the margins, and have those critiques stand in for reason.
"At sunrise the sun is only 'half there' ... there fore there is no 'day and night' just a blur! Day and night are the same thing!"
The training data used is part of all of this is a separate but related question.
So from my perspective, it's doubtful that this is the moat. Besides, for example, Claude Code in particular is so buggy (and always has been).
He talks about this in another recent essay https://stratechery.com/2026/anthropics-safety-superpower/
> If you own the user touchpoint, then you have meaningful lock-in, and the best way to own the user touchpoint is to be the canvas for everything they need to do. This, by extension, means that the frontier labs are on a collision course with software companies: it’s software that owns the user touchpoint, and it’s in the frontier labs’ long-term interest to not simply be a commodity input into software but to simply replace software outright.
true for anthropic, not true for openai.
For this reason alone I would also argue that the idea about an agent harness being sticky is a non-starter long-term.
i also expect you'll see markedly different results if you constrain yourself to small models. there even trivial harness improvements like Codex's /goal feature, and more capable basic tooling (e.g. semantic code grep, js-capable `fetch` tooling) make or break the actual task success rate.
The concept of “commodity” as defined above is a model, a simplified abstract representation of reality, but that does not match the reality perfectly (the map != the territory).
The author claims that a token isn't literally an ideal commodity, but neither is oil or wheat, many factors influence their real value (intrinsic properties, location, available storage at production, expected delivery date, etc.) so that no two gallons of oil in different contracts have the same price.
Is treating “tokens” as a commodity a worse model than treating oil this way? It depends who you ask! I'm pretty sure that a chemist working at a refinery would be more happy to see tokens being felt with like a commodity by his company than if they started viewing crude oil like one.
(Overall, there's way too much economism in that post, and way too few facts, and as a result the argument makes very little sense, the author basically wrote that both OpenAI and Anthropic are drowning in cash right now because compute scarcity means the price must be significantly higher than the marginal cost…)
Edit to add: I think what he's saying is more like "tokens aren't the interesting commodity, 'intelligence' is", which makes more sense. To carry on my gas and electricity analogy, I would say the same thing about gas being the less interesting commodity than electricity, because electricity can be used for a broader set of useful things. But both things are commodities, despite one being an input and the other being an output in this case, and the conversion efficiency is one very important consideration, but not the only one.
The value of intelligence is that it can solve my specific problems in ways that are satisfying to me. The example the article uses is a CRUD app - but the CRUD app I need isn't fungible with the CRUD app you need! It's not fungible at all, it's a specific solution to a specific problem that may have zero value to anyone else, and certainly cannot be replaced with anyone else's solution to their problems.
If we're comparing electricity to gasoline, then models are cars, and "intelligence" (I disagree that this is what LLMs produce, but whatever) is distance traveled.
Distance traveled is not a commodity. It's the desired outcome.
In your analogy, the "CRUD app" is analogous to the car, but that's not what Thompson is defining to be the unit of "intelligence". He's saying that some number of units of "intelligence" are necessary to create that CRUD app, somewhat analogously to how some number of units of energy are necessary to create a car.
But I agree with you that this concept of a "unit of intelligence" that he's using is probably too abstract to ever be usefully well defined.
* Me, as an individual, because I might not be able to pay price hikes, because my revenue (salary) is much lower than what they want and I can't support my expenses via huge bank loans.
* Again, me as a new entrant to the industry, LLMs are basically pay-to-play games, again related to price hikes, new entrants might not be able to afford paying those prices 24/7 - which you need when learning new things.
* Any non-US company, US can block the models which can disrupt the whole business.
* Even some US companies, for example if you operate in EU and EU somewhat changes their mind and follow the ICC and require you to stop working with Netanyahu (war criminal as per ICC), then following laws in EU, might create trouble to your whole business.
And read your data, see CLOUD act, PATRIOT act etc. etc.
No longer a theoretical risk in today's US political environment.
It’s not a new thing. Industrial espionage has always been a thing as well. So has bribery (for deals) been a thing especially by euro concerns.
Well sure, except with closed US LLMs you're basically just handing them data on a plate, and paying for the privilege. ;)
Very American really ... monetising industrial espionage.
"should" is doing a lot of heavy lifting there.
I agree it is hard to escape some form of legitimate need for intel.
The concern with intel and the present US administration comes on the checks, balances and controls side.
We are after all dealing with an administration happy to conduct much of its most sensitive business on Signal using off-the-shelf phones.
Some of it I think is selective picking. Similar to election interference where we know quite a few foreign states like interfering with our elections but we typically only hear about one of those states as being the culprit. Now, obviously we like interfering as well (Ukraine in ‘14 and Iran today, though Iran would be less controversial) and many others over the years.
The default email provider for most people in the west is Gmail.
Ah, cultural nuances. The title "Who's Afraid Of Chinese Models" is a riff on "Who's Afraid Of Virginia Woolf" which itself is a play on the song "Who's Afraid Of The Big Bad Wolf".
The title essentially means that the chinese models are being portrayed as the big bad wolf; but are they really the threat or are american frontier labs afraid of competition and commoditization?
It's also somewhat ironic because the author says that there is something to be feared -that western innovation will become dependent on chinese models, especially for cyber, if the american ones are restricted or unavailable.
DeepSeek, Kimi, Xiaomi Mimo, Qwen, Minimax, GLM, Hy3 and Ernie are always available, and I can't be happier
US labs have consistently demonstrated their intention to paywall higher intelligence. Eventually the paywall for “hyper-intelligence” will set a bar so high that average small businesses simply won’t be able to afford the bill of what is used by the top corporations to keep themselves at the top. That’s already starting, when it comes to the volumes of tokens top corporations are burning.
This is a feature of the system corporations want to establish and OAI/Anthropic are happy to oblige. 10% of a trillion dollar company is the same as 10% of 1,000,000 million dollar small businesses. Whose hitch would they rather ride, and which size customer easier to obtain to meet their revenue goal?
Not to mention, it would certainly be possible for the EU and other world powers to equally disincentivize use of US labs as a data security risk, since our top models are impossible to run in private lab environments without specialized agreements and there no access to the model weights for auditing. As best as I can tell, AI regulation is a dangerous game that is a hair away from isolationism.
In my opinion, Google is one of the few hopes in this area. There is still a paywall, but I feel like they are the closest thing to a Chinese lab we have (for frontier) in terms of their targets (real business use cases) and they actually have both the infra and already have a pipeline for small businesses into their products; they already have the wide non-AI customer base to leverage unlike Anthropic and OpenAI whose only product requires convincing people to use their (more expensive) AI.
* All modern AI is a perfect front for harvesting material for processing by NSA/GCHQ.
Given the criminal US' 5-eyes/9-eyes apparatus' atrocious war crimes and human rights records, this is reason enough to eschew American AI 'products'.
I'll use the AI created by the culture that lifts a billion people out of poverty first, not that from the culture that murders children every 15 minutes and lies to itself about it ..
China’s got plenty of blood on its hands, pretending otherwise is silly. America does, too. It’s quite easy for me to condemn both their governments and trust neither.
And only nationalist/racists ignore the fact that China has lifted a billion people out of poverty during the same period that the USA and its allies have destroyed countless other sovereign nations and left 36 million refugees from their illegal wars for the world to deal with...
China has a lot, lot better record on human rights than any member state of the criminal 5-eyes/9-eyes countries, which engage in massive human rights violations at atrocious scales every minute of the day.
1. US frontier lab unit economics are better 2. US frontier labs are moving up the stack making tools that are "stickiness" and will prevent users from switching.
For 1...he doesn't provide any evidence for US lab unit economics being better...the major input to unit economics is electricity...which is cheaper in China. And building data centers and connecting them to electricity is both cheaper and an order of magnitude faster in China. The main input that US labs might have an advantage in is in cost/access to chips, but that given the level of chip investment in China it seems unlikely to hold.
For 2...there's little evidence these tools are sticky. At least in programming, the trend seems to be tools like opencode that support multiple models and providers.
And even when they are sort of sticky, as we know on hacker news, people figure out how to point the tools they like to competing models even when the app doesn't official support it.
And every improvement in model capability makes it increasingly easier to make your own tools.
Wrote more on this in a blog post that has an earlier HN discussion: https://news.ycombinator.com/item?id=48982061
Direct link: https://larrysalibra.com/ben-thompson-is-wrong-us-frontier-l...
What is the cost of AI? The single largest ingredient is Nvidia profit margin.
Huawei accelerators are not as efficiency yet, but they don’t nearly extract as much margin.
Why would future revenue stay with the labs given this situation? This whole thing had an airline industry sized red flag on it that makes investing into frontier lab about as sexy as investing in United.
Maybe the token economy is some kind of reverberation of the airline reward miles economy, the emergency hatch to be able to survive under maximal supplier extraction (Nvidia is just the top of a monopoly stack here, even if they replace those chips, the HBM, ASML, Foundry layer can get their dues)
Sorta yes, sorta no.
A single 5090 consumes 450W - at Californian energy prices of $0.38 per kWh that's $0.17 per hour. And the card itself costs $4100 on amazon. So after 2.75 years running at full power 24/7 you'll have spent more on electricity than on the card. I would have thought most data centres being built today would have a design life longer than 3 years.
Of course you can throttle the cards to ~300W without losing too much performance. But also you need more than a single 24GB card to run most modern LLMs.
> 132kW
> ~$3.7-4M
So about 300x the 5090's power but 1000x the price. Roughly 9 years for electricity to exceed price at $0.38 and datacenters will show up in areas with cheaper power than CA.
LLM's aren't very latency sensitive and can therefore move to wherever power is cheapest.
Right now that's places next to aluminium smelters (which also like very cheap electricity 90+% of the time).
What matters most is $/completed task. It does seem like OpenAI and Anthropic are winning here even with worse electricity rates. Perhaps it is made up by the efficiency of Nvidia and Broadcom chips, which China can’t get in mass.
I do think that OpenAI and Anthropic are moving up in stickiness. My company has rallied around Claude. We are customizing Claude Code, adding knowledge bases for non technical people, writing skills for them, using Claude features company wide. It’s hard to move.
Meanwhile, I personally use ChatGPT outside of work. The memory, ease of use, habit keeps my subscribed.
However with the latest models Fable, Kimi K3, 5.6, it's getting to a point where I sometimes forget what model I am on without noticing a difference. And once I realize it because something may not be exactly like I expected it I won't switch for that work either because I don't want to invalidate the cache.
For the next work I will do there is maybe a 50/50 chance to remember to switch the model before I start.
That's not what I would call stickiness towards a certain provider.
Which is to say, this isn't really a lock-in/ stickiness vector (unless maybe the wording itself of a skill is hyper-optimized for a specific model)
Across sectors, China added 543 GW of energy in 2025. Next year, USA is expected to add between 70 and 80 GW of energy
AI is really all about electricity. AI could be completely fake and yield zero value whatsoever and the US would do exactly what it is doing now because the AI bubble is what creates the market for building new electrical generation capacity, which is needed for re-industrialization. Also why our friends in UK/Europe/China are so busy pushing anti-AI propaganda to try to undermine this.
You need to understand that the US added 40.3 GW of energy capacity in 2023. 70 to 80 represents a dramatic growth (nearly doubling) in new capacity.
Also China is just getting started. 8 out of the 12 nuclear power stations opening worldwide in 2027 are Chinese https://world-nuclear.org/information-library/current-and-fu...
If AI is "all about electricity" as you say then the US was barely ever in the running in the first place. There's no possible way the US could compete any time in the next 2 decades (putting aside a dramatic shakeup like war)
> Also why our friends in UK/Europe/China are so busy pushing anti-AI propaganda to try to undermine this.
Lol people make up the funniest theories when a political idea they don't like is gaining popularity. At least you didn't blame Russian bots
They do not want to life in the world that is and thats going to be, but in the past and the world they green ideology promised. Reality denial be a addictive poison.
AI is not fake and it does work, but what I am saying is that from a pure systemic analysis perspective, you can do the numbers, and even if AI was complete fugazi, the benefits you get from the electrical generation capacity, and the ability to fund it through private markets, which bypasses Congress, and locks in commercial contracts (often with foreign governments) which will be almost impossible politically to reverse, would still make it optimal from a strategic perspective. That is my calculation, and to the extent that it is correct, I would assume that the US Military's strategic planning apparatus would arrive at the same conclusion.
AI compute has some unique characteristics that make it especially useful for grid management. Moving consumer compute to the cloud means that the electrical use of that compute can be centrally managed. In an emergency, you can cut electrical use for consumer AI by 50% or more, because chips run more efficiently at lower power, and you can shift workloads onto quantized models, reduce resolution for video output, etc, to reduce compute, which leads to minor service degradation but not interruption. AI datacenters are also adding massive amounts of battery storage capacity, which is an additional grid buffer. For every GW in capacity added by hyperscalers that is creating a dispatchable reserve capacity of 50% under completely normal circumstances (hyperscalers do this internally to optimize their own costs) and then that number goes up depending on the scale and duration of the emergency.
That's not generally true, since there is generally still much reliance on NVIDIA. The true low cost providers are Google with their TPU and vertically optimized stack, and Amazon with Trainium. However, Google does not have their own frontier model, and Anthropic (who are partially served by Amazon) are also paying a premium for extra NVIDIA-based capacity from SpaceX, maybe soon from Meta too.
I don't know how the economics of domestic Chinese Huawei-based clouds (no NVIDIA) compares to the west, but since serving cost is mostly hardware depreciation and to a lesser extent electricity, they are not necessarily at a disadvantage (Ascend 950 costs roughly 50% of an NVIDIA H100), and more to the point it is irrelevant when considering US commercial use that is more likely to be using Chinese open weights models from US providers served on NVIDIA based hardware.
I think the real significance of Chinese frontier models being open weight is that it takes development cost amortization out of the US-based serving cost, while the US AI labs can't afford to do this. The US labs therefore need to reduce development spending to remain price competitive. The Chinese companies are of course still making money from the Chinese market, whether by selling API access or by other business models such as Ziphu making 75% of it's total revenue by selling services to Chinese customers who are running their models on-prem due to the Chinese apparently being very concerned about data privacy.
> I highly doubt that Chinese models are cheaper to serve on a marginal cost basis, they just seem cheaper because Anthropic and OpenAI are so supply constrained that they are charging far more than they would if there were sufficient supply to meet the demand for intelligence. [emphasis mine]
I guess I'm missing the part of this article where they bring hard numbers in to back up the argument here. What work was attempted? https://cursor.com/evals shows the previous generation of open models (Kimi K2.7) trading blows with the others, cost effectively. Composer 2.5 is itself a fine-tune of K2.7, and it's apparently quite token efficient, so why would it be impossible for a Chinese lab to achieve something similar? GLM 5.2 Max is also ranked above the lower end OpenAI models and is not far off in price.
It's weird to have this entire discussion about tokenomics without mention of the circular financing and debt raised by labs in the West, which can then essentially give away their capacity to end users. OpenAI giving away quota resets to subscribers like candy on Halloween while their compute partner Oracle's bonds is reevaluated to be one grade above junk? How?
I don't think you can make an argument about the future one way or another by arguing using the listed prices. The math is not internally consistent enough for it.
Because you're comparing retail price whereas the parent commenter (and the article) is talking about marginal (ie. inference) costs. American labs are providing a premium product and they're charging accordingly. Meanwhile for chinese models they're open weight so they're limited to how much they can charge without competitors undercutting them.
If we use tokens as a rough proxy of inference costs (rough approximation, I know) and look at artifical analysis benchmarks, you see that all the open models are behind the pareto frontier in terms of efficiency.
But if we have to look at what we think margins might look like, DeepSeek continues to host v4 Flash at the existing price despite competitors beating it in price (https://openrouter.ai/deepseek/deepseek-v4-flash), so there's at least one example of a Chinese lab charging a predetermined price despite competition. And no one but Moonshot is hosting Kimi K3 yet (https://openrouter.ai/moonshotai/kimi-k3). Perhaps there's room in the market for those who release their models to make margin on them.
And I believe my Composer example speaks for itself. The open models are behind but there's tangible proof they can be tuned for pareto frontier efficiency. See "Cost per Task" at https://artificialanalysis.ai/agents/coding-agents.
their competitors are discounted at around 33%, so it's safe to say that's the margin, maybe less if their competitors have worse caching or quantization. Meanwhile claude code/codex resellers selling tokens for 90% off API price, presumably by reselling usage from fixed consumption plans, which gives an idea on how fat the american labs' margins are.
>And I believe my Composer example speaks for itself. The open models are behind but there's tangible proof they can be tuned for pareto frontier efficiency. See "Cost per Task" at https://artificialanalysis.ai/agents/coding-agents.
But composer is a closed model? If it's really that easy to get better coding performance, why haven't the chinese labs replicated it? And this is all assuming the performance boost is real and not from benchmaxxing. Moreover if you apply the "street price" discount I mentioned above, American labs look far more favorable.
I look at that and think that they must be losing money hand over fist on something like this, not that this shows what their margins are like. If their margins are like this then I don't see why they'd be raising money and shuffling it around in circles.
> If it's really that easy to get better coding performance, why haven't the chinese labs replicated it?
Nobody said it would be easy! I just think it's possible, and that presumably they will get around to doing it at some point.
For an example of my real token usage for a day with DS: Input (Cache hit) 530,949,760, Input (Cache miss) 7,875,004, Output 1,389,685 - it is still 1/5th the price of Baidu (the cheapest) and 1/7th-1/10th the price of US hosts.
Also, Deepseek platform is not the same thing as Deepseek open weights. There is a major misconception that the existence of an open weights model means that it is the same thing as the proprietary platform offering, but that is definitely not the case.
Eventually we will hit a "good enough for cheap enough" and frontier models will hit diminishing returns (if they haven't already for a lot of types of work)
Don't think the rest of the world will sit on their hands while the US soaks up chips either, demand gets filled and if the US won't fill global demand for chips that's an opportunity to undercut again.
The other thing the rest of the world doesn't have to fund is the ridiculous valuations on these companies.
Unless you think the US can stay ahead just with model efficiencies, and that no one else will eventually match them, you are looking at the writing on the wall.
All that to say, the rest of the world is more than willing to eat your lunch, they have a dozen good reasons to, and they're already showing good results.
Just on the economics side, we've been here before too, US companies typically export their commoditization and live on brand royalties. Think all the cheap manufactured goods, the US doesn't make any of it. That's because the US can't compete on margins for numerous reasons, it's too expensive, I don't think AI is any different here except that the brands are currently valued in the trillions and I suspect that greed will be their undoing.
This can change quickly though, so it's not that big of a deal. If AI energy demands push the Gov to deregulate/fast track new plants, or the industry decides to build out their own generation renewables.
[1] https://www.iea.org/reports/electricity-2026/prices
Texas is the only is State that doesn't participate in the national grid (90% anyhow). Their prices are lower because they dont abide by the regulations that they would be required to if they did. Like weather proofing infrastructure.
This is also why their grid crumbles every time it gets cold, or hot, or... Tuesday.
See the 2021 grid failure for a particularly bad example. In that instance somewhere between 250-700 Texans died to keep their prices low.[0]
[0] https://en.wikipedia.org/wiki/2021_Texas_power_crisis
If I had to wager why, I'd say it's due to embracing solar on massive scales recently. Only a few years ago the US was competitively priced.
[0] https://www.globalpetrolprices.com/compare_countries/USA/Chi...
> Models are not free. Downloading them is free. Running them is not. This has manufacturing economics, not software economics; the idea that they are "free" is an economic category error as it relates to their actual use
I think a large part of manufacturing economics is illiquid overhead and the cost of expertise to set up and run your manufacturing line. Compute economics don’t have the same illiquidity nor do they require the same expertise or even specialized infra (current temporary chip shortage aside).
The implications of this are small players (e.g. your uncle running an inference server out of his garage) have comparably efficient marginal costs as big players. Compare this to actual manufacturing where small players have essentially no access to the manufacturing facilities of the big players.
Additionally, big players with a lot of compute who are not meaningfully in inference today (e.g. Amazon) have a fairly straightforward glide path to utilizing that compute to compete.
> This is because US labs are leading on cost efficacy of inference ($/task)
It’s possible, but I would need to see better data on this.
>A big part of training now is optimizing token efficiency. It's hard to distill token efficiency; that is perhaps why Chinese LLMs are so inefficient.
I think it’s fair to assume this is true, but also token efficiency is not a meaningful competitive moat. It’s not like these are secrets the Chinese will never figure out, it’s a fairly active research space and the outcomes are quantifiable.
This is where Oracles "datacenters for rent to run your own Chinese models" strategy will benefit. The LLM SaaS game is a lost one thanks to China.
I appreciate the transparency in explicitly stating their motivation for writing the article (a response to what the author saw as an overreaction to Chinese models), but I feel the article goes too far the other direction, with multiple unsupported leaps of logic, and overstating the stickiness of AI client products.
[edit]
This is my observation from using it without an specific context engineering to optimize for Deepseek's cache compression and sparse attention mechanisms. I am pretty sure that if you specifically structure your context to align to the cache compression boundaries you can significantly improve performance in the full 1M context, but there is not much reason to do this, because if you design your outer loop to work with shorter contexts that solution is portable and more efficient, so I haven't bothered with a optimizing for DS at this point.
Switch to a modern sampler like min_p or ideally a better one like top-n-sigma (it’s in llamacpp) and your “my model gets stupid at long context problems” will basically go away.
Unfortunately this fact is still not well appreciated yet despite nearly every modern sampling technique getting an oral wherever they get presented. Min-K just got an oral at ACL 2026, for a hyper recent example of this. There’s a reason they keep getting orals.
The field massively ignored sampling for mostly safety reasons and now the whole field incorrectly believes long context doesn’t work on small models. Long context is an out-of-distribution problem. Your sampler configured properly keeps you in distribution.
Oh and this is doubly true for quantized models. I run my qwen 3.6 27b with 4bit quants from unsloth and get excellent performance because my sampler stack is good and not the garbage that is top_p and top_k. Also, yes, you need to ignore the trash recommended sampler settings from the Chinese labs (they’re wrong/bad).
Is this really different from traditional software? Downloading postgres is free. Running it is not. You either buy hardware and assume the costs of owning and running that, or you pay to run it in the cloud.
IE, a lower param OpenAI/Anthropic model can compete with a higher param open source model.
So even if you are an American company who downloaded Chinese models in hopes of saving in cost, you still have to beat OpenAI and Anthropic in $/task which is very tough to do over the long run.
There being squeezed by their own stock pumping and SpaceX pretending to be an AI company is driving down the exit strategy. I’m guessing one starts going full Theranos and begins claiming full AGI or gets the US government to government cheese then hard. It’s going to be a few crazy months.
Over time, this compounds. More profits means more investments. More market share means more control.
As soon as manufacturing starts building this stuff more, it will commoditize. The hardware prices won’t be terribly larger than the original. We’ll have a “Bambu labs” style company to make the AI OS, whatever that is.
Regarding Price to Earnings... I'm not sure everyone fully fleshed out the end game for the frontier companies. It appears the pitch is that this current phase is a stepping stone to Artificial General Intelligence or Super Intelligence. I can understand the perspective of investors though... If you can get half of the world on these products at some point you will find something you can sell them even if its not the core product (i.e. loss leader).
This is a flatly false statement for most things powering backend applications. The AI consumer "doing real work" model, either for analysis, chat, or coding could well be more cost effective with closed frontier models.
But most of these internal glue business SaaS applications where engineers are integrating are not those tasks. It is those tasks which 1) drive immense amount of domain-specific data into the platform over time, and 2) are most encouraging of driving open model independence with no vendor lock-in.
Anyone on this site who has actually used ML models (more accurate in many cases) knows there's a lot of kludge that simply does not need a 5 minute agentic feedback loop to solve the problem. And they were solvable a year ago with lower class models. The token economics are exceptional and the anecdotes of a16z saying 80% of startups are productionizing open models is only surprising to people who think running your company on OracleDB in 2026 is a sound engineering decision.
The labs are not interested in the small, fast, single purpose end of the market. Google increased their pricing on Flash so much that it stopped becoming a cheap model; instead, they released Gemma 4 open source, which is actually easier to use from a third-party inference provider than from Google.
From a total token volume perspective, these "utility" models (classifiers, simple summarizers, small OCR models) will absolutely drive enormous volumes of tokens, at low prices and margin and modest overall market size. Because the models are small and the performance requirements are modest, and because their use cases are specialized rather than general, there are poor economies of scale: they can run cost effectively on rented small GPUs, and a big player doesn't get a structural cost advantage. These models are usually 1b - 30b in size, and can run on a rented 5090. I've productized these myself: I run millions of pages through a fine tuned 1b OCR language model that runs on 5090s at a cost far lower than commercial providers.
But that's not the segment of the market where GLM 5.2, Kimi 3, etc., play. They compete with frontier capabilities, and they are not particularly cheaper than OpenAI models at a cost per task. (I do actually think they compete well with Anthropic, because Anthropic's model efficiencies are poor compared to OpenAI.) And although this part of the market may not be the bulk of the token volume, it is the bulk of the market value.
That's because a lot of human knowledge work is too generalized and fuzzy for dedicated, fine-tuned models, so they are almost entirely different markets that don't particularly compete with each other. (Though if SaaS companies successfully build around verticals that can use small models applied against well-defined jobs, there may be opportunity to push the small/big capability boundary to subsume marginally more valuable tasks that today would require mid-grade reasoning.)
It may not be false but may be a "category error" [0]. Reserved GPU pricing & bulk inference pricing is 3x to 6x cheaper than "API rates", but renting your own GPU cluster (in this crunch) to run a 600b+ open weights is going to be "more expensive than GPT 5.6".
Even then, it remains to be seen if Huawei will pull their weight (and match up to Nvidia) as spectacularly as their fellow Chinese AI Labs have. If so, the WAICO alliance is ready to go all-in.
[0] Ben, and probably other "influencers" in this space, may be prone (knowingly or unknowingly) to favour points that make their conclusion for them (https://en.wikipedia.org/wiki/Motivated_reasoning).
But much like Ben's point that commoditization is a relatively novel concept to many in tech, it's not the consumer AI applications at risk of commoditization. They have distribution there.
It's the literally millions of engineers who are updating codebases with tools replacing workers partially or wholly. It's the supply-side where there's compression, and no need for distribution.
I would argue, given the enormity of the existing SaaS stack and how it integrates with the human machinery of personnel, that's where volume is. And that is clearly cheaper and a home run.
Commoditizing a ~$100B AI consumer market is no small feat. Commoditizing 20% of the $500B SaaS market, to say nothing of the underlying systems in the who-knows-how-many trillions "Big Tech" market (you're obligated to say that like the Kool Aid man), is shocking.
This is the story for Nvidia/AMD or cloud providers rather than OpenAI.
> With increasing inference as % of total compute, if labs create efficient models -- which they can, because they can create highly optimized models amortized over very high inference loads -- they can be low cost producers, and be competitive at $/task rates
It seems like there would be problems with this on both ends.
For general purpose models, everybody is trying to make them efficient, so you can't win just by being slightly more efficient. You would have to be so much more efficient that you can charge high margins while still capturing the majority of the market so that the high margins get multiplied by the majority of users and the users you leave on the table aren't funding open competitors. Meanwhile everyone else is also trying to improve efficiency, so one misstep and you're behind.
Example of where this can be a problem: You spend a preposterous amount of money to create an efficient model, then someone else publishes a paper with a new technique that gets a similar but incompatible efficiency improvement out of a model that costs a lot less to create. You have now spent an enormous amount of money in exchange for no competitive advantage.
And from the other end, one of the best ways to get efficiency is through specialization. A general purpose model can generate code or summarize a meeting transcript, but a special purpose model can do it as well or better with far fewer parameters and resources. But then you don't have a situation where one huge AI company has The Most Efficient Model, you instead have dozens of specialized models produced by independent sources that are each the best in a given niche. Any proportion of which could have open weights, or have an arbitrarily small advantage over the ones that are.
Moreover, these problems combine: Both the computing hardware vendors and the AI companies want the margin on doing inference, but the more of it one of them gets, the less the other does. If the AI companies were actually getting huge margins then it would be in the interests of Nvidia, AMD, Apple, Intel et al to fund efficient open weight models in the same way they fund Linux. Commoditize your complement. And those models don't even have to be better, as long as they're good enough that the closed models can't charge a significant premium and the margin shifts back to paying for hardware.
I'll agree that GPT 5.6 may well be the best given the above contstraint, but for run-of-the-mill dev tasks (real ones, not benchmark ones), GLM 5.2 still blows every other model out of the water.
Cost per task as a metric is a bit ridiculous because there are so many types of tasks. GPT-5.6 can do some tasks GLM could only dream of, but GLM can do some tasks 100x cheaper and better than GPT-5.6.
I could see AI ending up the same way where the customer captures most of the value rather than the companies. Open weight models are what make that kind of competition possible.
I notice that the article, and this discussion, hasn't mentioned or considered local models.
We can already run a low-spec model on a laptop. Because there is demand for this, it will improve and we will get better laptops and better local models. We will also see models being run on dedicated local hardware and called from the laptop.
If I can download a reasonably capable model to my own hardware and run it without paying anyone for either the model or the inference tokens (effectively making models and intelligence actually free once the hardware is bought) how are the Frontier AI Labs going to make any money at all, let alone enough to support their vast valuations?
Anthropic’s API pricing is getting impossible to justify. Anthropic previously had the highest quality models, and used their position to charge premium prices, enjoying inference margins of over 70% [0]. They could charge these prices because no other model came close.
But over the past month, the market has shifted dramatically. Over every single performance tier, Anthropic is being squeezed on price.
* Low end: DeepSeek V4 Flash runs at ($0.02/task), Xiaomi's MiMo-V2.5-Pro at ($0.03), and Haiku at ($0.24). Anthropic is ~10x more expensive than the Chinese open-weight options.
* Mid tier: Claude Sonnet 5 ($1.53/task) is nearly 50% more expensive than GPT-5.6 Sol ($1.04), nearly 2x the cost of GPT-5.6 Terra ($0.82), and 3x the cost of GLM-5.2 Max ($0.47). There is basically no reason to ever use Sonnet 5, the competitors are significantly cheaper.
* High end: Opus 4.8 ($1.80/task) and Fable 5 ($2.75) are the two most expensive models, and GPT-5.6 Sol ($1.04) and Kimi K3 ($0.95) offer comparable performance for significantly less. Less the fact that Kimi K3 will get ~10x cheaper once its weights are released and served on neoclouds with Nvidia hardware [1].
OpenAI priced their latest GPT-5.6 models cheaply in order to regain market share. When Anthropic clearly had the best models, their 70%+ inference margins were defensible. But today they are the most expensive option in every single tier. Unless they make significant price cuts soon, they run a serious risk of bleeding market share.
[0] https://www.mindstudio.ai/blog/anthropic-inference-margins-7...
[1] "American companies such as Modal, Fireworks, and Baseten will be able to serve Kimi K3, at one-tenth the cost of their Chinese competitors because they have access to advanced Nvidia hardware" https://x.com/rohanpaul_ai/status/2079027313455550839
Using a Chinese LLM will not put a Marxist under your bed.
Did you know Gemini is shockingly bad at French poetry? Hasn’t stopped me for using it for all other tasks though.
Interesting tangent - it might be taught to introduce stealthy backdoors in your company though. Maybe even across multiple PRs where each session puts a small chink in the armor, and together they allow unlimited access to the attacker who knows about them.
After all, LLMs are mostly black boxes. How comfortable would you be running a Chinese compiler?
They’re going to try their best to offload these investments into our pensions before the inevitable crash.
https://finance.yahoo.com/markets/stocks/articles/goldman-sa...
Too many people here buy high and sell low.
When Goldman makes statement, half the time it is prepped by an associate or two that has minimal experience and goes through a MD that enjoys the gloom and doom. That’s why they publish. Goldman makes money on both sides of a trade.
Correct. These chinese labs has proven that having just the model is not a moat, and the safety concerns were all just attempts at regulatory capture.
This is why labs like OpenAI and Anthropic are panicking and are racing to the exit before their valuations start being questioned.
Its the user base (with ads and upselling) and proprietary wrappers which will make money for typical customer.
Even enterprise customers arent going to be spending a lot on tokens. Once labs no longer have to subsidize trainings tokens costs will drop 10x and once models get burned on chips costs will drop 10x more and you physically won't be able to burn significant number of tokens unless you're deliberately trying to.
https://epoch.ai/data/ai-data-centers
I don't need 2.4T to do that; I'm doing it with 35B or 27B. If they get me a model in ~80B with a A5B or A7B, that will be the end point.
It's bizarre people, by themselves, believe all these parameters are getting them much more.
Lets be serious: if we as a civilization really wanted the advancements promised, we'd find the 1000 best scientists and give them free access to these models while the rest of us get personal GPUs for specific use cases.
But instead, we have to endeour this penis measuring contest for the infinite bikeshedding of the universe.
https://github.com/JustVugg/colibri
But its also a 800B sized model running on a ram constrained system with no GPU.
Techniques, GPUs, more ram, and faster disks can always speed it up. But the point being is they run on low end machines now. Its now an optimization problem, not a possibility assessment.
for VCs, breaking even is losing
(And that's good, that's how it's supposed to work. See eg prices for transistors or hard disks or solar power over the last few decades.)
In the end what companies pay for is not tokens but results. A DIY kit of models, mac minis or whatever, and a bunch of poorly integrated OSS tools doesn't solve their problem. For the same reason, people use Office 365 rather than running Libreoffice. And for the same reason things like AWS dominate the market rather than people DIYing their infrastructure together themselves. Most of the money is in polished turn key solutions. Which is what Anthropic and OpenAI offer.
The juicy market here is the enterprise market. That's mostly business users, not programmers. They'll be hooking up all their SAAS tools (which they also over pay for), and other stuff. They'll be paying for boring things like data residency, compliance, etc. And they need access to reliable infrastructure to run all this stuff. They'll want this shit to just work and not to be dealing with a lot of poorly integrated stuff.
Most of the billions invested are being sunk into infrastructure, chip design, and access to resources (land, water, energy) needed to run data centers. A handful of companies now own most of that infrastructure and they also happen to have the top models, researchers, the best tools, and warm customer relations. And they sell access via very convenient subscriptions with high enough limits that people don't have to worry about things like token cost. The game here is recurring revenue from customers that like predictable pricing, reliable quality of service, and iron clad compliance and data security & residency, and quality guarantees. These companies don't want to be chasing model quality and have to upgrade their entire company every few weeks. They want continuity and predictability. Mostly they just pay Anthropic, OpenAI, MS, or Google to take care of this for them. There might be some niche EU players that become a bit bigger. But I don't see a large scale switching to Chinese suppliers for a full polished alternative. The Chinese might give away their models. But I don't think they'll be generating a lot of revenue.
And if you want to run your own models, you'll still need infrastructure to run it. These four companies together with the usual cloud giants control most of that and as well of the supply of resources (chips, data centers, energy, etc.) in the EU and US markets. There's going to be a long tail of self hosted and gobbled together stuff but it's going to be a much rougher experience for end users and it won't likely be most of the market any time soon.
What is stopping China from gaining a majority market share, then, in terms of serving inference?
AI Sovereignty -- yes
Cybersecurity concerns -- yes
Latency -- no, unlike previous emerging IT workload types , inference does not have strong latency requirements. eg 1s of additional network latency doesn't matter to a 15 min, 10-turn agent session.
Cost -- ultimately this comes down to a nations ability to plug chips into warm shells. which forks into geopolitical / trade on the chips side and energy scalability and modularity on the warm-shell side. Even if you call geopolitical / trade a toss-up, China has the US beat HANDILY on the energy front, yearly they are deploying 10x power to their grid relative to the US, which is shooting itself in the foot at every possible moment.
IMHO chip tech will travel across borders, absent a breakthrough in analog inference, energy scalability will ultimately dominate.
There are also half a dozen other companies from China continuously hammering our clients’ websites.
I was wondering, what's in that cold dessert? Low and behold satellite imaging shows massive datacenter build outs, very cheap solar energy.
Few months ago something happened and the Geo location on data on those IP now shows "Shanghai" or "Shenzhen". A way to cover tracks? But mapping latency still points to fact that nodes behind these IPs are still operating around Xinjaing region
credit:
'You Can't Cheat Time: Finding foes and yourself with latency trilateration' https://youtu.be/_iAffzWxexA HN user: lopoc
Shenzhen vs Xinxiang is hard to do using this technique but Shanghai vs Xinxiang does show difference.
Assuming that China only distills is a huge mistake.
It’s no longer some backward place that does low value copying. Look at companies like ByteDance and Xiaomi.
Chinese companies aren’t just distilling, they’re acquiring data in the same way American companies did by paying people and crawling the internet.
The way I understand it, China has a few large companies that crawl the web at a rapid rate and build corpora. The government essentially wants select few companies to do this and then make the data available to other strategic companies operating within China.
Then there are data aggregators that buy data from apps, websites, and services, as well as systems like OpenRouter or Cursor, where companies can learn from the “traces” of coding agents, chats, and so on.
This massively reduces costs, as smaller companies like DeepSeek don’t have to do their own crawling or acquire data from 100s of websites and coding agents etc....
There are also companies in China that buy American LLM APIs and proxy them to companies within China. So, there could be 10,000+ companies using American AI products, while China logs all of this, understands how they’re being used, and trains on their traces.
My experience has been quite the opposite. I was using Claude Code almost exclusively this winter/spring and swapped to Codex earlier this summer. It took no time whatsoever to switch. And before Claude Code, I was using Cursor. Same story.
[edit: Oh and there was also a brief interlude with Conductor, though I think they're more or less just serving the underlying Claude/Codex harness]
For companies, these decisions are very sticky. Companies go through a lot of red tape to get anything purchased and approved, then they discourage change because it's a lot of work.
So the product that gets a foothold in a company sticks for a long time.
Then a couple years later a sales person convinces an exec that they can save some money by switching, so the switching game begins. Not necessarily motivated by the better product, mostly the price. My wife's company keeps switching their tools out from under everyone every year or two. Just when they get everything stabilized and everyone familiar with the new tool, some new contract is signed that moves them all to some other company's suite.
Similarly, I have made no ground in arguing to try to get Codex at the company I work for, which got Claude Code a year ago and sees no reason to go through the whole process of setting up any alternatives when Claude Code already works and is at the frontier.
You can choose a selection of different models within it, but you're not using Codex or Claude Code.
You install MCP connectors, specific skills, work around model/harness quirks, set security boundaries etc.
It's a lot of work, and most people will never want to change it once they have it working.
They will have people who don't understand the distinction between visiting Claude.ai and downloading Claude Cowork.
They type the words "setup MCP" into Claude.ai and expect it to automate Excel on their machine.
There's a pretty big gap between the things we talk about here, and where the world is at.
its so distracting seeing these types of confision.
every plugin is already just multimodaling their targets.
Probably the only reasons I would seek change are economical.
“I don’t really have a strong preference between the two” is another way of saying “the product isn’t sticky”, which is another way of saying “this provider has very little room to increase margins”
There’s little difference between Coke and Pepsi and the barrier to switching is nil, yet clearly the products have stickiness. People have slight preferences and become familiar with the brand and then engagement becomes habitual.
The effects on margins are irrelevant to this.
The cost of me moving around these different AI models and harnesses was pretty much 0.
They communicate through my own harness, and it's working pretty well so far. claude code is being overtaken by codex however because I noticed lately the accuracy of the latter is the best.
And I'm saying this as someone working for American companies.
They were all in the public domain too.
They do not protect individuals no matter how much people want to think they do. AI has proven this.
I think to level the playing field all copyright, trademarks, and patents laws should be eliminated.
If I want to make a marvel movie, I should be allowed to and profit from it.
AI let the cat out of the bag and there is no going back. We need to let individuals profit just like corporations can from what is considered theft right now.
But nah, let's do the most radical, least thought out thing, and absolutely destroy small scale creators. I'm sure Amazon will be benevolent and continue to pay writers in your scenario.
What is with 2026 and just conceding civilization to the worst actors, and then adopting the worst tactics/thoughts/concepts?
Even this isn't quite right. Publishers unambiguously have been screwed.
What has happened is if you can give the political/investment classes enough upside opportunity in the entities doing the IP theft then you're golden, and they increasingly have no problem with even pretending to hide it.
There will be no one debating the great minds of the twentieth century because the corporations made it illegal for people to republish critical editions of any such works.
The greats were recirculated every twenty or thirty years for centuries.
The smaller authors will vanish into the nothing and Western Civilization's 20th century onward will vanish and be almost forgotten forever because of that mouse. There will be more ancient literature remembered than literature after the invention of the printing press because they made it illegal to share printed materials for nigh on two centuries if an author was young when they wrote it. There will be more ancient scrolls preserved for the future than books of 20th century philosophy and science because the latter is a crime.
Only the oligarchs who pirated the books will have a cultural memory. They have cursed our era to oblivion when it comes to intellectual property because they can generate billions now for the mouse's henchmen.
Everyone posting on Stack Overflow knew it would be public and free. Book authors put an explicit copyright notice, which includes derivations of their work.
Because it's not a widespread phenomenon. A few large tech cos laid off large swaths of people a handful of times. That's only happening in those large tech cos. Most cos are empowering their employees with AI as a tool, not a replacement, and they're not letting anyone go (unless they refuse to use this new tool).
Of course, they're not hiring as much either, since their current teams can accomplish more with AI as a tool. Maybe the AI companies could give job seekers a bit of a discount, but that would be abused to all hell without crazy administrative overhead costs, so why would they?
because they are profit maxing. the AI world would be a completely different place if it's not for the open models.
As to the argument is that (only?) the Chinese labs are training on my data, I find this almost comical given the amount of highly-personal data companies such as Meta and Google have been harvesting for decades.
Even if the founders didn't want that, eventually the upper ranks will fill with MBAs and the board with private equity and they will make it that way. Their bonuses are based on quarterly performance not customer experience
Long gone are the companies who served their communities for decades or centuries, providing a stable return to the owners, jobs for the workers and value to the customers.
I'd rather China have my data than America. China might do something with it one day but America built exploiting it into the business model (disclaimer that I'm not Uyghur or Taiwanese though)
The perception of capability varies greatly between task. For my needs for example sol xhigh consistently outperforms fable xhigh.
If you run a model on slower hardware are you getting more experience? Surely its a factor of model output reviewed and not human time.
Love this.
Hermes is a better coding tool IMO. I can't put my finger on why but it just feels better. Maybe being true yolo helps.
There's just no trust in a country that is digitally totalitarian and hostile towards its own people. Do people ever look at the full sized Tianamen Square photos? This is not even the photo of the many people on the ground who were killed by their government and it is still insane to look at.
https://www.reddit.com/r/pics/comments/dgua6k/the_full_tiana...
Are you referring to USA, China or EU here?
People largely can't protest here right now, and US citizens are being killed. People are being sent to work camps in countries where laws do not apply. And our leadership is, right now, priming the American people for when they reject the results of future elections. Not to mention everybody involved in the last attempted coup was pardoned - after we were told that our current leadership had nothing to do with the coup.
The state of the US is much more dire than most people are letting on. And it's understandable why. Nobody likes bad news, and we all like to believe things will be okay. All I know is I can't take my phone into the airport. I can't go out and protest without risking my life and freedom. I can't drive anywhere without my location being tracked and logged. And, if I get pulled over, I must comply with any order, no matter how unlawful, otherwise I risk being executed in the street.
Some of these things have been going on for a while, and some are new. But all are real.
This is unfalsifiable, so not a great claim. You can continue to claim this forever without having to prove it.
>All I know is I can't take my phone into the airport. I can't go out and protest without risking my life and freedom. I can't drive anywhere without my location being tracked and logged. And, if I get pulled over, I must comply with any order, no matter how unlawful, otherwise I risk being executed in the street.
Wild, wild exaggerations, but you will point to your unfalsifiable claim to justify them.
The VAST MAJORITY of people in the US who engage in the things that you claim that they cannot, do so, and with no consequence.
I am no fan of this administration, but I am even less a fan of outright exaggeration (which is even one of the defining hallmarks of Mr. Trump)
It was just recently ruled that your phone can be searched extensively at an airport, without a warrant. The only reasonable thing to do is simply not bring your phone.
Notice what I am claiming and what I have claimed. I am not claiming these things will happen. I am saying there is a risk. People have gotten executed by the state for not complying with obviously unlawful orders. It has happened many, many times. And before you argue: yes being shot by the police is "public execution by the state without a trial or conviction". That sounds harsh, but that doesn't make it inaccurate.
Every time you are pulled over, there is a risk. Any time you enter a public space, there is a risk that ICE, which is a domestic peace-keeping military force, will arrest you or kill you. There is a risk. There is a risk you will not receive due process. Every time you go to a protest, there is a risk your phone will surveilled, sometimes remotely. We know ICE uses 5G devices to surveil the the air waves. Every time you go to a protest, there is a risk you will be subjected to tear gas or rubber bullets.
All of these things are real, are proven, and have occurred on many occasions that we know of. There is really no dispute here, so you can try to dispute it, but if you do, then you are outing yourself as a dishonest person, and of course then naturally nobody will waste their time talking to you.
Now, you probably aren't happy about this. Neither am I. But emotions such as unhappiness do not magically override reality. Whether these things are happening or not and whether they are a risk is a separate question. A question with only one reasonable answer: yes.
Actually, not bringing your phone is eminently unreasonable, particularly to an airport. Sure, you could wait in a line for a paper boarding pass, but if everyone were "reasonable" and did this, what do you think it would do to airport lines? What would happen when a flight gets canceled, and everyone rushes to the desks for help?
Buddy, birth has a 100% fatality rate. Life is risk.
It’s not an argument, it’s just not.
There is an amount of acceptable risk, you’re right. How do we calculate it? By making sure we follow due process, the law, and we hold parties accountable.
ICE is allowed to execute Americans. Yes, allowed. Because they’ve done it many times, and face 0 accountability. So they are allowed to do that.
Now use your knowledge about humans. When humans are allowed to do something, what happens? If we let people get away with murder, what is the end result? This isn’t rocket science buddy. Put on your thinking cap.
Ditto for the police. When the police make mistakes and violate your rights, they face zero consequences. Your average McDonald’s cashier faces more consequences if they forget your god damn ketchup. So the police are allowed to violate your rights, yes they are. So they do, obviously.
I’ll admit: this conversation is a little bit frustrating to me. Because this is not a deep analysis or conspiracy. It’s just very simple facts and then some very basic incentive analysis. It’s disheartening that there are Americans going around not putting any thought into the power structures of entities they interact with. It’s dangerous, too. People like you are the best target ICE and the police could ask for. They want people who lollygag and don’t care about their phone getting tracked, aren’t concerned about tear gas and rubber bullets, and who believe their rights won’t be violated.
at some point the conversation has to go past this reflexive "USA uses tech abusively -> but look at how abusive China is with tech -> ..." back-and-forth to acknowledging that neither party is your savior -- and then (hopefully) acting upon and coordinating around that understanding.
In practical terms you could get US and Chinese models to review each other, right. Depends what your use case is. Coding is kinda not so bad it is reviewable and immutable/traceable per commit. An AI app that is like a psychologist or something may be more worrying.
Seemingly better than everywhere else in the world, including where I live (Canada). Whataboutism doesn't really work when you use the least bad option as an example.
US is less hostile in the day to day sense, but they could look at your cloud data at any time because say someone in your company protested and exercised their 1A right, and the government didn't like that, so that sort of thing tracks directly to secrecy and privacy guarantees I might want from models.
I personally don't worry for my sloperating but if I ran a large company's AI I might consider it. Esp if in Europe.
The lessons from steel, solar and EV needs to be learned by all lawmakers. You have to respect and learn from how China Government puts the system in place for complete industry takeover and they have been very good at it. The problem with AI is that democracies will be inherently slow in adopting AI, unless something changes in the system.
At minimum, every democratic Government (US, Europe, India) need to build long-term AI vision and execute that no matter which party comes to power. Additionally, be ruthless about protecting domestic labs. It can only be possible if the intelligence pricing by domestic labs per productive task is in the similar range as open-weights models. Right now, it is not the case, even if the article gives the example of Sol vs K3.
Protecting domestic labs means not bailout, but fast track to cheapest energy, fast track approval for data centers, enforce some guardrails so customers get to use the open weights models only hosted in the country by US (or Europe) businesses. Without these protections, it might be a slow death.
It doesn't have anything to do with the form of government, it has to do with the aims of the government.
Nobody can predict 5 year out. However, the country that can be ultra efficient by making their governance, health, manufacturing, military, etc AI-native will be far ahead in the game.
But that's not really even relevant to the debate. Insofar as software has made the world more efficient, it doesn't matter where it's written. That's the point of open source software, there's no tacit knowledge. When a country loses its nuclear engineering capacity that's dangerous because it takes a long time to rebuild. The only reason the Chinese are already competitive on LLMs but haven't managed to make a state-of-the-art jet engine is because the latter, unlike software, is difficult to copy and you don't need to worry about something you can copy.
And it's the only viable tool the US has left. It's reasonable they are freaking out.
> To that end, here’s an even more interesting question around distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here?
> In fact, this paradox is the solution. I believe that open weight models are good for innovation (and, per the above, I think that labs on the frontier will be fine), but it’s a problem to be dependent on China. The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation, for U.S. companies at a minimum. Stopping distillation — which is literally just querying the API — is nearly impossible; the U.S. should go the other way and lean into a new copyright policy that both indemnifies the labs and also guarantees that what they learned fuels further innovation for everyone else.
That would prevent the facebook strategy of sucking up MySpace users and then defending TOS that prevent other social media apps from doing the same to them.
This has an element of stochastic improvement so it's hard to predict but the chance of the U.S. "winning" this "race" is pretty bleak.
You see this all the time in communities that have internalized hierarchy as a "good", little kings of shit mountain vying for less and less at a higher and higher cost.
An astute Chinese analyst could reasonably forecast that they had little chance of controlling the AI market due to sovereign trust issues, but would also note that AIs are just software.
When the dust settles the US still won't have factories, and the real value of AI models is still going to be embodying them and getting them to do real, consumer facing work.
Perhaps the most striking thing about the AI boom is how quickly the US abandoned the veneer of local manufacturing in favor of more expensive buildings producing nothing you couldn't make anywhere else on the planet...from imported parts.
https://xcancel.com/deanwball/status/2078133895766114412#m
I’m struggling to understand this perspective. Is he using the words accelerationist/decelerationist in a sense other than the obvious one?
EDIT: I searched his twitter history and discovered that his argument is basically “if you drive down costs, then OpenAI will have less money to invest in development, slowing down the overall rate of AI progress.” IMO this take betrays an overwhelmingly stupid degree of exceptionalism, but I guess that’s what I’d expect from someone working at OpenAI.
that sounds great to me. anything that slows down the pace and gives us a chance to prepare for, at best, massive job displacement, and at worst, robots turning us all into paperclips.
if open models are good enough then it doesn't matter, if they aren't then there's probably a return available in investing there.
Every company wants that!
"I am personally surprised the Chinese state continues to allow the open sourcing of models this good, given potential risks" what risks?
I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). Confused what this means
Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models. Confused again.
One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a "public good" which will ultimately be provided by the state as a kind of "digital public infrastructure." This future strikes me as a dystopian hellscape, but I've never met an open-weight models advocate who doesn't ultimately concede this is where things end. I don't understand this at all.
Can someone in the know please use plain layman's terms to explain what this tweet is about?
I think it's referring to the belief that LLMs are not the path towards AGI, and that LLM's, while useful, are not going to have the impact that the American labs believe it will have.
The Silicon Valley people like this openai guy, high on their own supply, are convinced they are building some machine god that will either bring about the end of the human race or utopia, they therefore cannot understand why the Chinese (or any other normal person on earth) are not afraid of chatbots and have other things on their minds.
I mean, what an appalling way to live. To take themselves so seriously and simultaneously be terrified of what they're building. If we are really seeing the collapse of the bubble now, I wonder what's going to happen to these clowns when their AI god fails... what will they move on to next?
Writer seems to have no clue how IP actually functions in China
I assume they mean the risk of opening up "forbidden" knowledge to the masses without adequate control, which the CCP hasn't historically been known to do.
> I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). Confused what this means
Yann Lecun is a pioneer in the field of AI and Meta's former AI head. He is famously anti-LLM, and considers the entire technology a dead end to achieving human-level AI. The author is saying the CCP has similar views (that LLMs aren't going to get exponentially better/lead to AGI) which is leading them to not control these models as tightly as they otherwise would.
> Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models. Confused again.
"AI accelerationists" = people who want AI to progress. According to the author these people should not celebrate open models because open source = less commerical value in LLMs = less investment into the field (because how are companies going to get returns?), and this will ultimately lead to slower growth.
The last bit is about government controlling AI vs commercial companies. According to the author the former is a dystopian hellscape.
IMO even if you think his points make sense, his job title ("head of strategic futures @openai") means they should all be taken with a massive grain of salt.
This is of course a baseless assumption. Let's say China created GPT 3.5. Then I can guarantee you that Ben would say "Western frontier labs are at a disadvantage when gathering data, because they have to follow the terms of service of Western media, and Western copyright law". Which we now know wasn't true.
And sure, some will say "but Anthropic can more easily block this as it's a single point of failure". But it's doable to overcome this. Without being "state backed".
China has a billion+ people that their AI can "study". Plus due to China's political structure, their AI has access to everyone's chats, comments and sites, scraping everyting.
Here in the US, with 1/3 the population, the AI race was lost before it even began. Plus in the US, all companies and people are doing all they can to restrict AI from scraping sites and peoples chats.
So I believe, China will end up owing AI.
I think you hit the nail on the head - right there!
Sure, any model that is not at the frontier can use the frontier model to generate synthetic high quality training data, so this can reduce significantly the training costs.
But at the scale of OpenAI, Anthropic and Google, it is quite likely that the (raw) training cost is very high anymore. Here's a few heuristics:
1. All the hyperscalers see a huge demand for inference. They can't deploy datacenters quickly enough to satiate all the demand they see. But, it's is impossible for the inference demand to be constant throughout a day or a week. If you use the times when the demand is lower than the peak demand (which is almost all the time) to dedicate the spare compute capacity to training, then your the cost of training compute is zero.
2. It is likely that increasingly a higher cost of the "training" is actually setting the guardrails, which is essentially post-training. As we've seen, without proper guardrails, the US Government won't allow you to serve inference. Anthropic was hit directly, but OpenAI delayed their 5.6 release as well to make sure the US Government is ok. This part of the training cost can't be reduced easily by using synthetic data generated by other models.
3. The frontier labs are also investing more and more in building an ecosystem around their models.
I am not a frontier lab insider, but take a look at the jobs posted on the Anthropic career page [1]. There are 74 jobs in "AI Research and Engineering" and by my count at most 15-20 are related to pure model training (of pre-training or RL type), and the rest are post-training, safety and security, alignment, interpretability, productivity and lots and lots of other things.
[1] https://www.anthropic.com/careers/jobs
Probably bad leadership
Imagine a scenario (theoretically possible but increasingly unlikely) where a US court decides that using "pirated" copyright data to train models is illegal. Now the AI developer has invested hundreds of billions of capital into a thing that is declared illegal and has to be scrapped.
This risk affects existing megacorps more than "startups" like OpenAI and Anthropic (and Chinese companies), because the megacorps have much more to lose. They actually have the cash to pay damages if the flood of copyright claims arrive at the door. This will not only bomb their AI development, but also the rest of their established businesses as well.
And thus I strongly suspect legal issues are holding them back a bit. Megacorps want to win the AI race, but not to the extent they stake the rest of their established business, while the newer companies' only product is AI, so they have to go all in.
Notice for example how Meta's Llama performed much more poorly after they got smacked by a bunch of lawsuits claiming that they torrented a bunch of copyright data.
(Disclaimer: I'm an outsider and everything I base my speculations on is public knowledge.)
Of course the Chinese companies have incredibly talented researchers, and smaller, better organized org structures which account for the rest of the difference.
I remember some feature lauded by Gemini was reverse engineered by the open weights guys in < 30 days.
If they dont publish some technical information its hard to protect in the US, but conversely, once it is published smart people from outside the copyrightosphere can start working to reverse engineer it.
>3. The frontier labs are also investing more and more in building an ecosystem around their models.
Theres nothing there that isnt immediately replaceable.
Indeed. But that was not my point. My point is that we still have this old impression that training cost is dominated by compute and it is hugely expensive, and the Chinese labs can short circuit that by distilling the American frontier models. I don't think the training compute cost is a big factor anymore for the American frontier models, because of the reasons I gave. If the Chinese models can get the training compute cost down by a factor of 100, that's not going to make them 100 times cheaper, and not even cheaper by a factor of 2. Maybe 10% cheaper or so.
Distillation is a technical term with real meaning, and historically requires logits which Anthropic does not provide.
"Generated training data" is the correct term. It's not an "attack". And Anthropic undoubtedly also generates training data for each new generation of models, yet you never see them claim Fable is a distilled Opus.
2) The word "attack" is standard security vocabulary. Per RFC 4949:
There are hundreds of named "attacks".3) The "attack" part of "distillation attack" refers to distillers creating tens of thousands of fraudulent accounts, using proxies to bypass georestrictions, deepfaked IDs, and paying real people to pass biometric KYC checks. Who then blended this in with real user traffic to conceal their behavior.
It doesn't refer to the AI training technique in any way.
If they acquired this data without the fraud, you'd have a point.
In a way you could see this as a case of Robin Hood. The US companies exfiltrated all the data on the planet just to hoard it for themselves now and accuse anyone who tries to get a piece of that back from them, and the Chinese labs are distilling it to offer it for cheap.
Obviously a bit more complicated than that but it still holds pretty well.
Sure, and large-to-small is a key part of the definition, and why it's called distillation (cf concentrating something). When Anthopic use synthetic data generated by Opus to train Sonnet or Haiku, then this can correctly be considered a type of distillation.
When Anthropic accuse Chinese companies of "distillation", it seems they are using this word to refer to two potential uses of their model outputs:
1) Using Anthropic model outputs (aka synthetic data) as training data, especially for reasoning, for Chinese models. This really isn't distillation though, since (unlike when they distill their own models) Anthropic don't actually provide the reasoning in their model output, only a "summary" designed to hide the actual reasoning. You can't distill what you are not given!
2) Another way Chinese companies may be using US LLMs is for "LLM as judge" where you are just asking the model to use it's expertise to judge/rate something that you provided yourself (to provide RL training rewards), although for coding you really want hard rewards which are easy to obtain, not fuzzy "looks good to me" ones.
Of course Anthropic are trying to pull the drawbridge up after themselves and their TOS says you can't use their models to develop anything that competes with them, and this seems to be what they are generically referring to as "distillation" - any use of their models that they suspect is being used by the Chinese to improve their own models, not just what what might more technically be called distillation, unless you want to define that word so broadly that it does mean this!
And if just partial output is all that it takes to declare a model distilled, then every model trained on Internet content since 2023 is now technically a "distilled" ChatGPT and Claude model.
It's a post-facto attack, which doesn't sit right linguistically to me.
2) This is a stretch: it allows Anthropic to arbitrarily define "attack" via TOS, and ignores the fact that the generated training data is literally paid for by the "attackers".
"You're trying to kidnap what I've rightfully stolen."
Lol. Isn't this literally many of the same tactics OpenAI and Anthropic used to scrape the internet? So now it's an "attack", but previously it was just "training".
It's actually a very goated term but not everything is causal, it also has precise technical meanings (although those get blurred too given that causal can mean anything from intervention proper, to mere depdnence on something prior)
I like Anthropic, I don't think all their talk of safety is bluff and bluster, or at least, I want to believe that the people who left OpenAI because it had lost its focus of helping humanity still want that to be their main goal. However, yes, it seems that business fears are once again causing those in charge to turn "we want to help humanity" into "we are the only ones who can help humanity, and therefore we need to be the most profitable, and the only survivors".
If you want the former ideal to survive, at Anthropic and outside of it, you need to be willing to collaborate beyond profit incentives and recouping capex. Show other labs a commitment to research and community and they will follow. Better to bring teams together rather than implicitly say you distrust them, pushing them that way instead.
Whatever the strategic picture is at the top, at the bottom of the market "weights you can download and run on hardware you already own" is the whole ballgame, and right now that's mostly Alibaba's to lose.
That's a big claim that his whole thesis rests on but is largely not backed up. Where are the apples-to-apples tokens-to-answer benchmarks that he's using - doesn't look like there are any, just a handwavy implication that US models are more token efficient, which they may be. But how is there so little effort in establishing this point in the article? And US labs may be in much different situations from one another: it's known that some labs like OpenAI bought big, early on compute and may have secured better pricing.
His article also does not mention the average price of electricity in China vs the US, which it seems like China leads on, and probably has the political power to more heavily subsidize. While I agree the COGS is often overlooked by top line benchmarks on coding tasks, etc, it seems that he's running on a big assumption while claiming "labs on the frontier will be fine".
I have sonnet do the thinking, deepseek does all the tasks. I've massively reduced costs with this approach.
On one hand there's the relentless barage of American propaganda. I get that a militaristic society needs an enemy to fight against lest they turn on each other. I get that if you tell people a bad guy is coming for your jobs or your lives then you can maybe get your workers to accept worse conditions and living standards which increases profits. It's ghoulish but there's some logic.
On the other hand I can't see China doing anything except minding its own business. No tariffs. No bullying other nations. No wars started. No threatening allies. They don't let their people waste their lives on brain rot or gambling. They largely align with UN resolutions. They respect international institutions instead of always being the asterisk.
Has American propaganda just failed to work outside it's borders? It's not landing at all
> - Supplier A will sell 10 units of the commodity for $20, earning $10/unit
> - Supplier B will sell 10 units of the commodity for $20, earning $5/unit
> - Supplier C will sell 5 units of the commodity for $20, earning $0/unit
> ...
> Bankruptcy risk is where fixed costs come back to the forefront: Supplier C has both fixed costs (like potentially R&D spend) and also may have taken on debt [...] It can’t price its commodity with these costs in mind — remember, the market-clearing price approximates the marginal cost of the highest-cost unit needed to satisfy demand [...]
Why can't Supplier C price their fixed costs and debt into their product? The entire reason Suppliers A and B are earning $10 and $5 per unit, and not more, is because they cannot meet demand by themselves and are therefore at the mercy of how much Supplier C is willing to charge. Couldn't Supplier C just refuse to offer 5 units of the product at a price that would bankrupt them?
Sincerely, an interested observer of business/economics.
I don't know if I agreed totally with the assessment of the risk Chinese labs pose to US labs though, in particular I think the main part I wasn't sure about was this:
> I highly doubt that Chinese models are cheaper to serve on a marginal cost basis, they just seem cheaper because Anthropic and OpenAI are so supply constrained that they are charging far more than they would if there were sufficient supply to meet the demand for intelligence.
How true is this? My understanding from Deepseek's original paper was that they focused heavily on optimising training and inference costs, in particular so that they can operate on cheaper (and more accessible to China) hardware.
It's possible I'm just not in the loop, but nobody seems to talk about US models innovating in this way (I'm just talking about cost-to-serve/train, not saying US AI companies don't innovate in other ways).
It seems to me at least, like there's a fair bit of evidence that AI shifting to a price based commodity market (vs a "best-model takes all" type market) would put China at a significant advantage? And even more significantly, require a pretty hefty correction of company valuations in the US?
I also heavily disagree with this no-marginal cost in software distribution view whenever I see it, bit rot is real, and someone is paying a marginal cost whenever they do an update. You have to re-distribute with changes whenever anything changes. These costs are just hidden because things are ad-supported or bundled in some way. These costs are also kept low because of standards and open source, but could become high anytime. Additional licensing also has costs.
That said, I couldn't agree more with the last paragraph, charging a high price for models would be better than denying access for any model that wants to stay relevant.
Personally, I think models will increasingly become specialized in different areas, some good at X, others good at Y, and we might see workflows that mix multiple models.
Following this argument the key for each player will be the underlying cost structure and serving capacity to offset the upfront R&D cost.
The cost infrastructure will be driven by access to cheap electricity and cheap chips. The capacity will be driven primarily by depth of pockets now to buy all available supply in chips/mem/data center building capacity. While China is certainly in the lead on cheap energy, I am wondering if they can/want to beat the > 1tn USD being spent on data centers right now. Following the example in the article:
If company C from China sells 10 units for 20 USD produced for 10 USD they pocket 100 USD.
If company A from America can sell 100 units for 20 USD produced for 15 units, they pocket 500 USD or 5/6th of the market's profits.
I think we see this with Meta being paranoid about internal Claude usage, to avoid inadvertently distilling[1].
If distillation is a driver, then smaller American labs could be distilling, but are not for legal reasons.
But that's a big if we just don't know for sure.
1 - https://cryptobriefing.com/meta-restricts-claude-code-codex-...
Sure as an indie hacker, you could go download the weights for a Chinese model with a VPN, and then attempt to run it at home by building your own GPU cluster but these large models require quite expensive hardware to run on and so it makes it less likely than anyone would invest that much capital to do something that is illegal. There's no way for them to sell a legal service using those tokens. So it can only be strictly for personal use (the Govt won't care because very few people will have that kind of money and risk appetite). The other option will be that there will be some shady third-party providers in foreign countries who are willing to sell tokens from these models to US consumers knowingly.
So under a ban rest-of-world gets to use cheap open-weight models but American companies/individuals must only use only ‘approved models from US for-profits’? Doesn’t seem like that kind of protectionism will be popular or politically tenable. Not so long ago US chose cheap TVs over maintaining the country’s manufacturing base.
(Despite what you wrote it’s also really hard to imagine that enforcement wouldn’t leak like a sieve. Unser sufficient economic incentives [which are the predicate for the ban], loopholes will be found.)
Whether or not distillation matters a small amount or a big amount, still interesting:
https://www.whitehouse.gov/presidential-actions/2026/06/nati...
Source: https://martinalderson.com/posts/the-upcoming-ai-margin-coll...
It gets particularly hairy because models themselves can tune their "token verbosity" to manufacture demand for compute. If compute was such a precious resource, you'd think we'd be complaining that the output was too terse.
The ability for a vendor to determine ex post facto how much a query costs is a similarly new economic phenomenon to zero marginal cost.
This was my assumption as well. It's also generally true of 'traditional' deep learning models that inference cost is expensive compared to training.
But the cost per token for inference has been very quickly dropping. I don't recall where, but I recall about ~50x down from GPT3, even as model complexity has increased. Even with agentic systems, there are lots of optimization opportunities. I'm less assured about claims like this.
Is he casually assuming a singularity has already happened? A regular first-mover advantage I can understand, but those have been squandered or lost many times before.
The actual difference is how much scrutiny and time was put into the Mythos / Fable and GPT 5.6 release. Making it feel like “these are a big deal”. Spring and summer THAT was the AI story
Then Chinese labs release models that approach Fable performance. We’re shocked they just seemed to appear out of nowhere.
It’s less about the gap closing. It’s more about the weight we put into Fable-capable models.
Today chineese deliver that promise and usa people freak out like they have any skin in this game. Enjoy the ride leader of the free world....
it would be beneficial both for openai and world (most likely)
Also, the Hidden-Agent problem exists in every model, and is a persistent tangible risk independent of whatever team people cheer for at the games. Let us remember, every LLM nuked all of humanity 92% of the time in simulated war games. =3
I'm pretty sure the U.S. has done both over the last couple of years -but I'd love to be proven wrong. :)
China will just do a better job -- if they do this at all -- of storing, collating, indexing, and using data from their state sponsored and championed AI labs to use that against the US. [1]
The US under this admin is doing the same, attacking universities, allies, it's own citizens.
The two governments are operating more or less the same. Ergo the ai models from each country's ai companies shant be trusted either.
[1] https://www.abc.net.au/news/2019-05-16/grindr-why-is-the-us-... The late Lindsay Graham comes to mind. Even larger bombshells probably.
I see massive risks in belief the inferences drawn from strategic information cannot be seen. So if you depend on some position remaining inside a secure facility but you drove to it from data outside that secure facilty, The likelihood that an inference model can derive the same idea is very high. Collation over public data is not inherently secret because you used a secret model or secret weights.
A more simplistic take might be that the fear is not actually driven in the secrets, the fear is "the emperor has no clothes"
What if there's a way to extract the commodity of intelligence from smaller models?
I've seen for many use cases it's well enough. :)
I think it is the right move to protect American interests
Ben Thompson is wrong: US frontier labs are right to be panicking
https://news.ycombinator.com/item?id=48982061
If EU build some SOTA open source models, they will design a different story
I'm amazed that no one is talking about proposals that are surely being discussed in Washington and pushed by SV lobbyists to restrict Chinese models on national security grounds, or other some other basis.
The belief that Bytedance could engineer a finger on the algorithmic scales to serve the interests of the Chinese Communist Party led to a lot of debate in Washington, and ultimately resulted in TikTok being divested from its Chinese owners. Huawei is shut out from the U.S. market, which limits its business even in markets where it's not banned because it's effectively stamped with a scarlet letter.
IMHO, Chinese models are headed for a similar fate or at least a showdown in Washington or the courts because they are supported and/or controlled by entities which ultimately serve the CCP.
Is this an assertion that is backed by evidence?
From the Elon/OpenAI trial:
> On the stand in a California federal court on Thursday, Elon Musk was asked if xAI has used distillation techniques on OpenAI models to train Grok, and he asserted it was a general practice among AI companies. Asked if that meant “yes,” he said, “Partly.”
https://techcrunch.com/2026/04/30/elon-musk-testifies-that-x...
[1] : https://imgur.com/gallery/ai-models-on-atrocities-B7DKUXc
I just asked ChatGPT about the Vietnam war and it did not say that it was purely the US's fault: https://imgur.com/a/zmiOyuu
It also didn't seem to have a problem describing people killed for left wing ideology: https://imgur.com/a/LhH9saL
These are both with the free ChatGPT membership, as I do not have a paid membership anymore.
I know this is a common trope to bitch about, but honest question: did you actually try this before you commented?
ETA:
I was curious what something that was trained around me specifically (fairly typical lefty American progressive) would say, so I asked Claude (which I have a paid membership for and have discussed political things about many times). The answers were broadly similar: https://imgur.com/a/caFgKHH
how is running servers supposed to be 0 cost, while running ai inferrence isn't?
For a SaaS business, running servers isn't free. But compared to the cost of running GPUs for inference that you are selling, it almost is. The company I work for is a SaaS company. We have a single production server. A couple of QA servers. All hosted on Hetzner. Monthly cost for servers is less than $400. This generates a few million dollars a year in revenue.
If we were in the business of selling inference, our cost of providing the service, for the same amount of revenue would significantly higher.
Even large businesses like Microsoft, Meta, Google have operated with similar margins. Cost of running servers, compared to revenue was very low. But inference changed that, in a dramatic way.
A single response from kimi k3 requires hardware that cost between 500k and 1m dollars up front and draw over 20kW. Each request costs at least 5% to 10% of the charged cost.
Frontier labs that thought they could Rupoor[0] the entire creative class, transferring the coercion premium of copyright ownership from Hollywood to themselves. In their eyes, copyright should not apply to them, but also their models should have exactly the same value as a copyrighted work.
Stratechery also argues the US should explicitly make training fair use and forbid terms of service that prohibit distillation. I'm in support of the latter, but NOT the former, even though I normally hate copyright. My reasoning is primarily that copyright is one of the few legal paths available for a rando to go and put the work of an AI frontier lab in legal jeopardy. In the EU and Japan, such legal action has already been foreclosed by similar law. And while free distillation would obviously be preferable, it's also much more of a legal long-shot. Getting America to do anything that even smells like taking property away from the powerful is impossible[1] - it's our zeroth amendment. But we can at least hack the property laws that currently exist to cause problems for the frontier labs.
And, to be clear, if distillation is OK but training is not fair use, distillation is still OK. The output of an AI model is never copyrightable, because copyright only protects the human element. Essentially, this would say "don't train on humans, but absolutely rip off and steal the shit out of other AI labs and give it to the rest of us."
[0] In the Legend of Zelda series, Rupoor is anti-money - collecting it decreases the amount of money you own. I am using it to mean "turn someone's asset into a liability".
[1] Given that America was literally created to protect a wealthy land/slave owner class from disenfranchisement, either from above or below, and the last time we did this we literally had to fight a civil war against that same owner class that installed a new owner class that has largely remained today
Or any US hyperscaler with GPUs to spare can decide to serve the models for a reasonable cost/token.
You don't have to send China your data.
I have been working on a project with about a dozen generation tasks, each of which comes with a fixed token budget. The nature of this system requires that most tasks be completed by distinct model families.
As a result, I tested ~50 models across as many model families as I could gather, frontier and open weight, API (gateway and direct) and self-hosted. Evaluation was based on a set of cosine similarity validations that was repeated across ~50 different embedding models.
Interestingly, frontier models did worse on the tasks than open weight models. However, when it came to costs, the picture was reversed: frontier models were much, much more token-efficient. In fact, almost no open-weight model was able to meet the initial token budget, while almost all frontier models did. Moreover, open weight models struggled massively with reasoning, in terms of latency and token consumption.
I also found that the latest models did not perform better than older models. And any a priori benchmarking data was utterly useless.
So, I ended up using a set of open weight models without reasoning, as it turned out reasoning as well as frontier negatively correlated with the tasks. However, before I knew this, I had spent a lot of time running each available reasoning level for each model.
Lastly, as an aside, when it came to embedding models, size (dims as well as model size) did not correlate with quality, once a hurdle figure (~2k dims) was met. In fact, sweet spot was 3-5K, and for my (text-based) set of tasks, dense models tended to outperform MoE ones.
This might be a simplistic take, but my biggest worry with depending on Chinese models (and, by proxy, open-weights model development) is that the US can deem them a national security risk at basically any time, and Ant/OAI have minimal interest in making frontier-level models open-weights.
Regulated companies prohibit Chinese models in anticipation of the ban-hammer from the feds, so for data-sensitive work, they're stuck with LLaMa, gpt-oss and Gemma models (which are good and serve as a good-enough base for sft, but seemingly not as good or as expensive as Chinese models)
I suppose the USG can do the same thing that China is doing and bankroll/subsidize that effort; whether they will is for fate to decide.
Nonetheless, this article made it clear that nVIDIA is the real winner in all of this. Shovel selling to the extreme.
lmao
But as inference becomes cheaper, some of the market will move to self hosted inference. I look forward to someone supplying small servers designed to run inference locally.
With or even without open models these companies are selling compute, and we've been making that rent vs buy decision for 60 years.
>Right now defenders are effectively banned from using Fable or Sol for cybersecurity because of Trump administration directives; that means the best alternative is using models from a country which has been trying to weaken our cyber defenses for years. This is insane!
I understand some guardrails are needed, but it is becoming increasing problematic manage them without a strong public discussion.
You techbros need to get off your ass and go to work.
1. training new base models are expensive for sure, but fine-tuning them are relatively inexpensive enough the labs can continue to do so forever. the main reason why frontier models are so good is because the massive input they generated from user usage. they are using that information to strategically build better training data. and this is why no other models can catch up, til now that is. but if chinese models are good enough, and free to host, and cheaper to use, then the consequence is the frontier labs will lost valuable user inputs and the chinese labs will gain more. as time goes by this will be a domino effect.
2. nvidia is not only the player in the hardware scene. amd mi350p is getting popular, and huawei is pumping SuperPoDs. what does this mean for us? chinese models will surely use chinese hardware, and optimize for them. the other people will pick amd because compare to nvidia they are cheaper. with open weight models and open source inference stacks, they are freely to experiment and improve the stack, thus further lower the inference cost and nvidia dependency. and they even plan to build their own inference hardware, too. and nvidia loses market share meaning all the fund it gives to openai or anthropic will be cut, too.
and you say there is nothing to afraid?
We have got very far from Cicero's coining of the word 'intelligentia' (from inter legere, a 'reading between' and hence discernment) when people talk about 'intelligence' as a commodity
People have been decrying the 'cheapening' of the word intelligence for over a century now, going back to Psychology's adoption of the word and coining of nonsenses like "Intelligence Quotient". "Artificial Intelligence" is just the latest degradation of the original humanistic meaning, and now people aren't ever bothering to prepend 'artificial' to their idiotic use of the word
This from OpenAi's Head of Strategic Futures "Some observations on Kimi: It's a very good model! I don't think its performance can be explained away by distillation or anything like that"
https://x.com/deanwball/status/2078133895766114412
China's strategy of spending billions on training these models and open sourcing these models away is strategic - they want to kill the US LLM industry at any cost.
To win on the AI front by any means necessary.
Why is it when Anthropic and OpenAI spend billions trying to beat each other it is competition, but when the Chinese companies do it then it is trying to kill the US LLM industry at any cost.
The US federal government spends billions in subsidies via the US Chip Act, and bans chip sales to China to support US companies.
But the implication is that somehow Chinese competition is illegitimate because "strategic".
Why wouldn't it be? China is pumping out AI research and researchers at a staggering pace and there is no inherent reason why western models should be better
The United States' real advantage over China is freedom. Chinese LLMs simply can't compete with American ones when it comes to the humanities, creativity, entertainment, or financial transparency. As long as the U.S. continues monetizing these strengths, the compounding effect will make it virtually impossible for China to surpass the U.S. at the product level.
> monetizing these strengths
> real advantage over China is freedom
Please tell me if I'm unfairly paraphrasing but these seem to be your main argument and they seem to be oxymorons