The lesson of the last 50 years of the computer and software marketplace is that free and low-end eventually wins.
- PCs destroyed minicomputers. Mainframes survive, but serving a much tinier portion of the market than they used to.
- PC office productivity software destroyed expensive professional products.
- Windows (low end) and Linux (free) completely destroyed the UNIX marketplace, and again, have taken huge market share from the mainframe world.
Ignoring the huge Chinese open-weight models for a moment:
- The training costs and resource requirements for frontier models are unsustainable. The high price, and social pushback, mean that the American companies producing these models are precarious.
- There are enormous financial incentives for research results allowing for cheaper, less resource-intensive models of high quality.
- Local LLMs on consumer hardware are akin to the PC hobbyist world of the 70s and 80s.
Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now.
Getting back to the Chinese models: They allow for new competition against Anthropic and OpenAI, basically SaaS renting out these very capable AIs much cheaper. That will just accelerate trends.
I do think open-weights models are going to "win" in the sense that they're probably going to be dominant when the hardware to run them becomes affordable. (which might be a while). Although I guess you could probably rent the GPU's yourself to hypothetically save on costs. (I'm a little skeptical -- I've heard of companies doing this and the inference bills are surprisingly high -- assuming the sources are correct. I don't know if a lot of people really want to be advertising "oh god our bill is horrible")
I'm sort of baffled by what the entities that train the open-weights models get out of it though. Is it just a direct play to undercut the US providers because they view them as a threat? I just don't really understand the business model behind it.
If you're NVIDIA then open-weight models are a classic example of commoditizing your complements; cheaper models mean more people buying GPU's to run them. [1]
Nvidia selling more GPU's at the cost of it's datacenter business is pretty close to Kodak selling digital cameras at the cost of film.
The data center side is so bloated anything that eats into it is a huge negative. Their data center business brings in 20x the gpu market. Local open weight models will be what pops the bubble and China will do anything in it's power to enable that pop.
I wonder what the thinking inside NVIDIA is at the moment. They have countless examples to learn from here, about the danger of not being willing to cannibalize your high end products. But, of course, there’s a reason that there are lots of examples of this sort of failure.
There’s plenty of competition that would be happy to attack them from below, though…
> Nvidia selling more GPU's at the cost of it's datacenter business is pretty close to Kodak selling digital cameras at the cost of film.
This assumes a) AI is a zero-sum game, and b) we're actually talking about on-prem AI will replace cloud-based AI. I think neither statements are true.
AI is like compute: we'll need all sorts of it, in various sizes, everywhere. I'm sure Nvidia whats to own all the workloads.
On the other hand, I do think open weight, like open source, will win in general.
Hmm this sounds like an incumbent missing a paradigm shift because they didn't want it to disrupt their core (usually enterprise) business, although riding the shift would have ultimately delivered an order magnitude larger business.
Classical example is Microsoft actively undermining mobile because it threatened selling Windows or enterprise licenses.
Or Yahoo fighting Google's model because the latter model's didn't depend on taking enterprise deals to rank results.
The business model is making your country more innovative, which makes its citizens richer
Americans used to do that too, spectacularly.
Just consider the two alternatives: one is your industrial sector with this incredible new automation and analysis tool available for free. The other is one where it has to pay huge chunks of its resources to overseas companies.
> I just don't really understand the business model behind it
There’s a lot of value in the same sense there is a lot of value in controlling what Google search results are shown and what people see in the Twitter feed.
They get money from subscriptions and tokens, same as for closed-weight providers. Yes they'll lose some traffic to hosting services, but many users prefer to use the original training company since they have a guaranteed-correct implementation. Similar business model as open-source SaaS companies.
Some companies (most notably Deepseek) also manage to host their own LLMs so efficiently they undercut all third-party hosting services.
No one cares which company they use. They care about cost and does it act in a way they expect. Expecting anything else is pretty laughable and goes against human nature. You either create a moat so deep no one else can play or you have the government force users to use your products.
I think part of it is definitely to weaken US providers and the US economy as a whole. China has a completely different domestic economic structure and motivations from western cultures... it's probably closest to a fascist economy mixed with a Maoist cultural ideology behind it. There's definitely winners and losers and the state tends to have tight controls over everything though.
I also think the restrictions on OpenAI and Anthropic are somewhat short sighted. In that the guardrails dramatically limit efforts towards securing your own software in many ways. Yes, it's also "dangerous" and maybe there should be a means of identifying "domestic" or otherwise "secure" accounts for those allowed to use the models without the same guardrails in place.
Most of the growth is. If we removed AI, the US would be in a severe recession right now. Most of the absolute market cap, employment, capital, and other metrics that are not directly coupled to growth paint a better distributed picture.
Of course, growth like this can only continue for so long without changing that.
Yes, it damages its image, this is further made evident given the amount of propaganda that follows each time. Why would you invest in claude or codex if you just read how China's stuff is better?
GDP growth in the United States is in AI and healthcare.
AI capital expenditure is around 5% of total US GDP.
Housing right before the 2009 market collapse was around 6.7%.
Biggest issue I see is housing has more real value than AI expenditure. Demand is real and isn't based purely on a few companies valuation or marketing spin. Nearly 2 decades later we still haven't caught up to construction rates before the 2008 collapse. When the bubble pops it's going to really suck.
> I'm sort of baffled by what the entities that train the open-weights models get out of it though. Is it just a direct play to undercut the US providers because they view them as a threat? I just don't really understand the business model behind it.
In China, it's because they are being heavily subsidized to do the research activity. It's not really complicated -- if you allocate public money for people do to a thing, they will do it.
China wants the US economy to flounder. Building our entire growth model on software that can be copied and taken by a small group of people will have no possible consequences.
China is looking after their own interests, but they absolutely don't want their largest export market to struggle. The global economy is not a zero-sum game, and the idea that it might be is the root of many of our policy issues in the US.
It’s not that simple. If the US economy goes into a recession, it will take large sectors of the weak Chinese economy with it either directly or indirectly.
It’s probably more accurate to say they don’t want American LLMs to become dominant. The huge US data center build out doesn’t depend on Anthropic and OpenAI anyways. Those data centers can just as easily serve Qwen or GLM models.
It is largest customer for now, but they try to boost internal consumption as well as diversify client-base. US is also competitor, so China is interested US to fade at least in competitive areas.
I think lot of Americans overestimate how much China needs them. (Understandable, I suppose, since the UK made a similar mistake vis-a-vis the EU about 10 years back.)
From a Chinese perspective I expect China’s largest customer is China. These days it’s actually kind of wild how many Chinese consumer products aren’t (and won’t be) available at US retailers. And a lot of them are quite good.
For a while now China’s wanted to reduce its dependence on the US for a variety of reasons. And undermining the US tech industry, whose products the US government likes to use as a cudgel, may serve that goal quite nicely.
Assume, crazy business, that china wants what’s best for the world. Recognizing the danger of an AI arms race rapidly producing uncontrollable superintelligence, it focuses on open weights to reduce the economic incentives for further advancement of AI beyond the “highly useful for humans” stage.
Seems like the only thing that could avert an intelligence rapid take off. Everyone wins except for shareholders.
I don't think it is true that they necessarily want the US to struggle, I suspect it's more self interest.
LLMs seem to be one of the biggest innovations of the last few decades, China probably just wants to make sure it's not being left behind and/or made hugely reliant on the US for what seems to be turning into a piece of critical infrastructure.
China has an effective strangle hold on some key sectors (solar, rare earths) and I am sure they relish this position and the leverage it gives them. You'd be careful not to give away that same leverage to a competing power if you can invest a few billion now and cover your bases.
ICYMI: the entire US technology industry was built on government subsidies, and then Elon Musk hired a bunch of 20 year olds to destroy the US R&D pipeline because of gay mice or something.
Allocating public money to basic research is good. We should do more of it.
I don’t think they need the hardware to become affordable (as a regular end user).
They need their models to be good enough and cheap enough. Then the rest will follow. Companies will figure out how to host them for you efficiently, and you pay them monthly.
I don’t think I will ever want to set up a home server, no matter how inexpensive the hardware gets. At work I still use Cursor (with Anthropic models usually) because it’s paid by my employer, but for private stuff, I’m already using cheap models with OpenCode, and it’s extremely cheap and surprisingly capable.
I think there was around half a year between where the best models became good enough (last year December?) and where the cheap models became good enough (couple of months ago?).
China is the factory of the world. They don't need software to win. Rather they prefer software is free and they can win in hardware. So if AI inference is free, they can put it in as many hardware components as possible and sell them in the market - think toys, cars, tools with chips manufactured in china optimized for the use case. In long term you tend to commoditize hardware. We have thousands of device types of cheap x86, and with linux/bsd software on it coming from china. Why do we think GPUs will be different.
We just sacrificed our entire lucrative SaaS market to this, and perhaps even big tech itself. To the altar of AI.
And now it's all going to become commoditized.
Billion dollar software will be commodity. Salesforce. There are orgs already moving to their own internal tools.
It does not seem hard now to rebuilt Google Search, Google Chrome, Gmail, Gsuite, Netlify, Vercel, Cloudflare, Vimeo, Twilio, or even Stripe. The cost barrier has to have dropped 1000x, maybe 10000x.
We have millions of engineers with the talent to do this. Many of whom are unemployed and have savings and nothing better to do. They could easily carve these markets into pieces.
We shouldn't shut down open weights. It's too late. They'll win, and that's a good thing. Big tech was a thermodynamic bubble of high energy waiting on the dam to burst, and now it has. The genie won't go back into the bottle, and that's totally fine. It's progress.
Now we need to rebuild our factories and supply chains and energy and resource inputs. Because the back half of this revolution is going to be robotics and factory automation. If we don't have the connective tissue in place, we're really going to hurt.
We'll do well if we regrow manufacturing. If we don't, we might be in for a world of trouble.
> It does not seem hard now to rebuilt Google Search, Google Chrome, Gmail, Gsuite, Netlify, Vercel, Cloudflare, Vimeo, Twilio, or even Stripe. The cost barrier has to have dropped 1000x, maybe 10000x.
We just did this song and dance with the tariff war between US and China last year. It was revealed that 3% of China's exports are purchased by the US, and that they are fully economically prepared to take that to zero when push comes to shove.
China is a centralized economy that has been fighting a tariff war against the US for the last few years. If we see an economic collapse like 2008 China loses money but we lose a whole lot more.
China still ends up with production and a easier ability to integrate with other countries. We became a powerhouse after WW2 when all the other countries had their financial bases destroyed. We became a superpower providing the products to others to rebuild. Our market is all about short term growth, service economies, and goals driven by whoever is in office at the time. If they don't have a plan around this exact eventuality I would be beyond surprised.
Basically their system is more robust than ours to a major market event. Who pays for expensive services when they are choosing between food or keeping their economy going. People will pay for the equipment to keep their economy going. Who provides that.
> I'm sort of baffled by what the entities that train the open-weights models get out of it though
if you do the training then you're in control of the output. For example, recommending your products/services or failing to mention your competitors. You could also automatically introduce backdoors into code deemed interesting, i'm sure all governments are very interested in having that influence.
People are happy to pay $50/year per seat to have the extra features and to not have to deal with stuff.
There's an issue at the margins here:
$1000/employee is a massive cost - it has to be deeply justified.
$50/employee is like ... $2 out of your pocket. It's an incremental cost. The CFO is happy to pay it if there is a lot of value.
A lot of software is in that later category.
Imagine if gasoline was 1 cent per litre - and there was 'free gas' but it was a pain to use, and you had to check a bunch of things. You may just pay the 1 cent.
AI is not quite that yet, but these dynamics will play out eventually, for a lot of things.
I’m suspicious of some quotes here, “80% of startups using Chinese models,” doesn’t seem quite right to me. I just interviewed at several startups and they were all using the US models. Maybe they have some minor use of Chinese models but the bread-and-butter of most of these businesses model use is the Claude and Codex subscriptions.
I can believe it with startups as they are trying to keep costs down. Established Enterprise however, is a different story and is where the money is usually. If those startups become successful the story may change, but most of them will fail.
We're using deepseek with the idea that we would switch to something better when more of our customers are using the ai features but it ends up deepseek is awesome for what we're doing and so we probably won't switch because it's so much cheaper.
The ai libraries we use let us switch models with just a configuration change.
Similar story here. DS models are absurdly good value for mid-end tasks. I've found DSv4 Flash to be ~10% the cost of GPT-5.4-mini/Claude Haiku at similar performance.
We used to pay OpenAI >1m$/month for fraud classification, NER, etc. Sadly the US companies no longer care about non-coding-agent uses.
I imagine uptake will continue to increase as the corporate infra improves. Right now it's still bad - for example, AWS Bedrock is awful, models are months late and implemented with basic errors. Google Vertex is even worse. Finding a decent provider is the hardest part.
If you're self-hosting a model as a startup (e.g. using GPUs), you're almost certainly using a Chinese model. If your a startup outsourcing (e.g. using tokens), you're going to be using a US based model
Neither of you are clear on what exact location in the world you're sampling from here, could be you're both right, just missing that you're talking about different places.
It depends whether it's talking about using models in the product versus for development. For example I have a few apps that use open weight models, whether on-device or via API if Internet is available, but for development I use Claude, Codex, Cursor for Grok etc.
When developing AI services, Chinese models are cheaper. For the AI models used in actual services, like uploading an image and receiving a response, they use Chinese models.
On the other hand, when developers are developing, they mainly use US AI because the quality is better.
When developing AI related services, they prioritize Chinese models due to lower API costs.
It seems like the article didn't make this distinction.
So the claim that Chinese AI is the top choice for service level AI isn't entirely wrong.
What I’ve seen is coding is usually done with US frontier models and anything that is part of a feature on an app and runs at scale on the API is a Chinese model because they are dirt cheap.
It will cost you more than it saves to use smaller Chinese models to code; because of the repeated work. That has been slowly changing recently, but with much larger Chinese models, however those models are so expensive they're much more price-competitive iwth the US competition.
But for actually providing end-user AI features, particularly simpler ones, the US isn't even in contention. The costs and limitations just outright kill those features conceptually.
I believe this depends on country. I would believe, US startups trust US tools more. In other countries which see china in a positive light https://www.pewresearch.org/global/2026/07/15/people-in-many... I would expect them to use Chinese models due to their lower cost.
The point is that for many tasks today, and likely all tasks before long, that the open vs closed will not be a differentiator. There are many open models much better than gemini, yet people still use gemini.
It's like picking AWS vs GCP. Yes it is a business decision, but one that will not likely affect the outcome of the business.
> but one that will not likely affect the outcome of the business.
We don't know that, that's the point of my statement about changing the question.
Do successful companies opt for the US/Closed models? If they do or don't it's just a correlation but it means something. Maybe it's just causal of companies being able to get more funding because the ideas are better so they opt for the more expensive model (assuming it's better).
Hmm if you assume the 80% is uniformly distributed between successful and unsuccessful startups/companies, which by default you should, then yeah, your proposed statement is true.
Data would be needed to argue the 80% skews unsuccessful
We use US models for everything in practice, but we are looking at open-source models right now. So it may just be a turn of phrase hiding the reality. I can't imagine anywhere near 80% are relying on open-source as their primary models.
Cursor may make up the difference. A ton of companies use Cursor and Cursor's UI and billing model pushes their in-house "Composer 2.5" model pretty heavily - which is a modded Kimi K2.5 model under the hood. Anyone using Cursor is likely using Chinese models at least some of the time.
It depends, I guess if they still dont care about losing money or they already had subscriptions then they keep using them, likely either openai or anthropic.
If they noticed the bill going up, like with github copilot, they are looking at the alternatives.
It’s a sneaky statistic. You could say that 100% of the startups I worked at used Windows laptops because at least one person had a Windows computer somewhere.
If you saw the engineers you’d see 80% Macs and 20% Linux laptops.
The statistic would technically be true.
I use Chinese open weight models a lot, but they’re not what I reach for when I’m doing important coding work.
I don't use any open-weights models for coding, but I use them heavily for document categorization and extraction. At millions-of-documents scale even the smallest models from OpenAI or Google cost more than running a small model on my own hardware, and for a lot of tasks I don't need the extra intelligence of proprietary models.
Yeah. People should absolutely be _trying_ the Chinese models, and experimenting with running things locally, but the noise in development is genuinely all Claude and Codex.
I put my foot in the mobile comparison the other day, and will again. If you were to go back and be a mobile dev in 2010 by all means specialize on one platform, but play with both as a professional interest to stay realistic. Here it's important people have access to US/Chinese/Other, open/closed, local/cloud and that this remains. Don't become a blind Claude guy or a open weights fanatic: that way lies disappointment.
Frankly I don't see much difference between Claude and DeepSeek except in the price. I use the Pro plan to work for a customer of mine and I topped up $2 on DeepSeek in late May for personal use. I worked on 3 projects and I still have $0.72 left. The Chinese companies will win on price, not openness.
Dunno I put to deep seek the question what I’d need in hardware to get full service k3 with lower latency in western pa (I was making like it was a business proposition.) It says several million dollars at minimum.
I'm using a ten dollar a month US model to vibe code startup ideas. All my previous startup ideas i had to hire a graphic designer and back-ender or two to help. I use to be a web design front end enigeer since 2009 yet those skills are dumb now, so now Im a vibe coder.
The model I use to vibe code with I am just going back and forth with. Since Im building it as I go using an agent doesn't make sense but I guess that's where all the token usage comes from? Pardon ramping up my skills via vibe coding this one idea for about a month and have never hit any quota and or have gotten anywhere near my limit.
the distinction may be between using the coding agents vs using models for products. for example where i work we're talking about dropping opus for a chinese model for the in-app agent (which is very expensive to run)
> When entrepreneurs walk into the offices of Andreessen Horowitz (a16z), a big American venture-capital firm, the odds these days are that their startups are using AI models made in China. “I’d say 80% chance [they are] using a Chinese open-source model,” says Martin Casado, a partner at a16z.
This is very different from what the author portrays. It may be the case that many pre-funded startups are using Chinese open-source models (somewhere in their workflow). But what percent of startups that survive more than a year (either with funding or revenue) are still doing this?
I imagine their pitch is: "look at how well we're doing using open source Chinese models! We'll do even better once we raise money to be able to afford frontier models!"
The way the author presents this quote makes me think he had a preferred narrative and found quotes to back it up. Or he's just a very uncareful reader.
I don't really think the author is misrepresenting the quote in any way, it just seems like many people are reading it as a hard statistic instead of a reference to a properly attributed quote.
Just anecdotally though, my company is not a startup, well established and well known and has already started investigating, purely for dev purposes (not product), using Chinese models - this was spurred by costs rising much faster than expected.
So while I agree that I don't think it's anywhere near the 80% level across the board - I wouldn't be surprised if it starts moving that way.
Found the narrative, it was right there at the bottom. Now I remember reading some of his other stuff and it's very much in the same vein.
> I care about having open technology that can be run in the public interest, aligned with the public’s values. Threads like public AI, federated services, and open research have traction but need backing. Getting there in the US needs more nuanced strategy and support than we’re seeing today.
It's also important to note that several PE firms have signed contracts with model trainers to specifically use their model in their owned companies. I know Anthropic signed a deal worth hundreds of millions to acquire users just a few months ago.
Could it be that the startups that embed LLMs in a product will prefer Openweights for a more stable economic model, while SWE who use LLMs as a tool prefer SOTA to produce code?
There may be confusion between 'product tech' and 'development tooling'. It's entirely plausible most of the startups that are A: building AI tech, B: early stage, and C: raising from top Sand Hill Road VCs, are currently prototyping their product tech concepts starting from open weight models.
A VC partner meeting with early-stage founders is focused on the viability, uniqueness and defensibility of the IP tech stack not what tooling the coders are using. The developers could be using Claude or GPT 5.6 to develop a tech stack based on open weight models.
This quote sounds like it’s about companies building products on AI to serve a tweaked or harnessed version of that to others not the developers in those companies model of choice for coding.
makes it sound like the second part is a continuation of the first quote from the same source, but actually the second link is just some random person’s substack post from almost a year ago.
Not everything runs on paid models. Claude and Codex are frontier models, but some people have much higher usage needs and finite budgets that force them to self-host. And if you're self-hosting, you're very likely running a Chinese model
Our small team (~6 devs) is still using Claude Code because we're still on the cost-per-seat-month plan. If we were being pressed to pay per token, we'd be re-evaluating for sure.
Came here to post the same thing. I've noticed a lot of Chinese model astroturfing on HN over the past 60-90 days or so. Many upvoted posts in all conversations about AI touting how great the Chinese models are even when performance isn't the topic of discussion.
it's artificial general intelligence, not advanced. the point is that they'll be smart across the board at some point, super-intelligence is a whole separate issue.
I’m obviously a tiny, insignificant data point as a solo freelance dev, but I watch this space closely and I canceled my Claude Code subscription today.
I don’t think they meant exclusively Chinese models. Many companies including the one I work for uses big US model for most of the work but data sensitive ones are on prem open source ones.
They probably use Claude and Codex for their actual development, but for the products they actually build and deliver to customers I imagine a lot use open-weight models.
If you're putting a lot of your money and time into a business, do you really want it built on a service only hosted by one company that will turn it off eventually and you have no recourse?
If you build something against an open model you can take that and run it anywhere. If your favorite model provider stops hosting it, you can go elsewhere, you can go rent GPU instances, you can even shell out and buy hardware to run it yourself if you've got the capital and it makes economic sense. Change some API keys, update a URL in your config, and you move on.
If the government decides that proprietary model is too good and so it gets shut off, what do you do? If a proprietary provider decides it's not worth it for them to continue hosting that model, what do you do? If that provider silently updates the proprietary model and it makes your app broken, what do you do?
This is a very strange article considering that Llama, the mother of all open-weight models, has led to anything but success for Meta.
Also, enterprises don't give a rip if models are open. They care about zero data retention (and sticking with whatever vendor they're already using).
This blog post is suspiciously close to being a restatement of what Alex Karp recently said on CNBC[0]. It's important to remember he's the CEO of Palantir and hardly a neutral observer.
There are many reasons to celebrate open models, I run them myself. However there's not yet enough evidence that 1. America is losing the AI race (pardon jingo-ey phraseology) and 2. American AI labs are losing because their models are not open-weight.
There has always been room for both closed and open source software.
In this context, open source/self-host means do it yourself, closed source means you trust someone else to do it for you.
In the long run open source always wins because of the community effort, customization, network effects and price.
Internet protocols are open, anyone can setup a website, host their own email server..., but most people don't do that, they rely on someone else to do it for them.
People still pay for Windows rather than use Linux because most people and companies have better things to do.
The main threat to American AI companies is not that consumers will self-host, but it's that new hosting companies will appear that will host open source models and offer them (cheaper) to consumers. It'll be like the web hosting market before the era of cloud computing.
> In this context, open source/self-host means do it yourself, closed source means you trust someone else to do it for you.
I don't think that's what it means in this context. Hardly any users are training their own models, after all. And I don't think Windows/Linux is a good analogy - OSes have inherent platform lockin that LLMs don't.
Really the only 3 things anyone cares about are capability, cost and data privacy. Sure, the cost axis for open weight models needs to include the cost for hosting things yourself, but I think the bigger reason that businesses have been throwing billions at Anthropic's and OpenAI's models is that they have had the best models, and they've had big releases every few months. Their biggest Achilles heel is that if/when their improvements start to plateau, open models may catch up and businesses will start to scrutinize their AI spend a lot more.
> considering that Llama, the mother of all open-weight models, has led to anything but success for Meta.
This is also a strange way of framing it, though. Llama was released as a research project, it was never intended to create some vast ARR revenue stream or reframe the way people look at AI. If Meta wanted to exploit it for personal success then they had lots of opportunities to do so.
With OpenAI and Anthropic's profitability under question, it is up in the air whether or not America's stance towards AI will work. If they can't convince the world that they're a proper software business, then China's philosophy will win by-default.
> it was never intended to create some vast ARR revenue stream or reframe the way people look at AI. If Meta wanted to exploit it for personal success then they had lots of opportunities to do so.
They release the base model as open source, everyone uses it. They make a paid version, no one uses it.
They make no money from the open source version, they get no social credit from it. Where is the benefit to having an open source model?
> Where is the benefit to having an open source model?
Research. Llama is and was a research project, intended for researchers. You could make this same critique of Microsoft's Phi model, Apple's OpenELM or OpenAI's OSS. None of them were intended to be kingkillers, all of them are experimental in nature.
You might not have followed the space at the time, but there was a real race to implement the transformer architecture with fewer overall parameters than GPT-2 and GPT-3. Llama was revolutionary for sticking the landing without being entirely lobotomized, the "benefit" was that the model was usable on a local machine. Contemporary projects like Flan-T5 and GPT-J/GPT-Neo were entirely displaced, Meta's AI mindshare went ballistic for a few months and probably propped up billions in exit liquidity for executives and former employees.
It seems to me that the US AI industry has bet the farm on the idea that AI will enhance AI itself, so any small advantage will magnify recursively into an unstoppable advantage. Therefore it is vital that they spend as much as is necessary to be the first to that small advantage.
At the moment, I can't say that I see this happening. It's hard to know whether it may happen in the future.
From what I see it seems like we're hitting the top of a sigmoid curve in the model's utility for coding assistants. Going from "it does 90% of the job" to "it does 94%" of the job is a legitimate improvement, but it's not a phase change, and it's probably not worth paying multiples more for. And coding assistants have turned out to be the killer app for AI; it still isn't really working out in a lot of the rest of the industries of the world.
At any moment, theoretically, someone could find some new way of making AIs that breaks this sigmoid and propels us into a new one. But that's not a great thing to bet the farm on.
I'm not sure I'd say "China" wins if this particular strand of American AI fails. Falling back to an open weights model and making money on the serving of the models wouldn't take all that much economic realignment for the US and would be the natural outcome of any sort of fire sale of current AI assets. However the devastating effect on the stock market if the market comes to the conclusion that this current round of AI can't be profitable without falling back to such an economic posture and some more years of people adjusting to it can hardly be overstated.
2-3 years ago MAIR was on a roll with Llama 1 2 3, Zuck was on his rehab tour to be a cool guy, and Meta as a whole was pumping record numbers after record numbers. I can't believe how that falters so quickly after the addition of Alexdandr Wang.
Don't think it is down to Wang or MSL but Meta's focus on "personal AI" led them to whatever strategy (OAI missed the developer market too, which Ant then captured; leading to several high profile departures at OAI, coincidentally hired from Meta). The original Llama team themselves started Mistral which hasn't gone anywhere. The simple fact of the matter here is, Chinese firms have the money and the talent to rival the US ones in this field, should they as much miss a beat.
I am no fan of Wang but he came after Llama got caught benchmaxxing Llama 4 rather than training a good model. My read is that Zuckerberg tried to buy his way out of the problem like he always does, and he ended up overpaying for a lemon.
At the time the whole thing was led by Yann LeCun who seemed to spend more time arguing with people on Twitter than figuring out new techniques to make Llama the best. Meanwhile Deepseek was figuring out large scale RL on kneecapped hardware like H800s and how to scale architectures an order of magnitude bigger with MoE.
The behind the scenes element you're not aware of is Anthropic is going to companies reliant on their models and demanding HUGE one time fees (100 million+) to continue using their models or they will be cut off. This has happened to several larger companies I and others are invested in.
This resulted almost every time in "screw off we'll train our own models or use refined open source ones instead" leading to a lot of anger at Anthropic by CEOs these days.
We all want an alternative and Anthropic and OpenAI need to charge more than they are worth to pay back their investors and everyones stuck now.
Pretty sure he’s talking about Cursor and other ai coding startups. There was a lot of drama between the two in the past year and iirc anthropic cut their capacity.
I mean isn't the explanation simply that llama was never good enough, even when it was released? I hear (no data) lots of people using gemma4, at least a month or two ago.
Isn’t it basically impossible to run the newer high quality Chinese models locally, even for a corporation? The better they get, the more they need a data center. So, the ‘better’ Chinese AI gets , the more it will just be a service run on Chinese hardware competing with ‘our’ lower latency AIs .
The open source character of the models is irrelevant if you need a nuclear powered data center for inference. In the end it is just another internet service.
> The better they get, the more they need a data center.
That's true.
> So, the ‘better’ Chinese AI gets , the more it will just be a service run on Chinese hardware competing with ‘our’ lower latency AIs .
This is not. Them being open models means that any hosting provider in the world can host them as well. You get to pick and choose the provider the same way you'd pick and choose where to run a Linux server.
The cost of inference is not insurmoutable for many corporations who may run a small datacenter out of their headquaters or branch offices.
Larger models absolutely have a higher barrier of entry, but its a cost under a few hundred thousand as opposed to the millions necessary for a datacenter built as core revenue generating infrastructure.
Cost of inference is small when compared to the cost of training new models. Which is the real advantage these Chinese models have. Someone else has already spent the capital needed to create the model.
With a reasonable upfront investment and a few trained staff, its very possible to run these larger Chinese models in a well managed fashion.
The real calculus is if this up-front investment and associated lifecycle costs are over or under the costs a corporation may simply wish to dump into a cloud managed service like OpenAI.
> Also, enterprises don't give a rip if models are open.
They care about control. I see many of my enterprise (or just-below-enterprise) clients very annoyed at OpenAI and Google after 2-3 years of model toil, where they had to constantly re-calibrate onto new models, on tight externally mandated deadlines, with little certainity. Now they are reaching for open weight models instead, that they currently host with the same inference providers, but have the option to in-house if push comes to shove.
I do not understand the logic going into these companies. Flagrantly violate all IP in Human history, essentially claiming domain over the heritage of Humanity... And... Try to privatize it? When the technology -- and data -- are both public domain to begin with?
It is ming-boggling stupidity. If there is talk of bailouts as the dust settles, there it would just be further evidence the system is ethically, financially, and intellectually bankrupt.
Look up the English Enclosure acts. They brought about a large-scale robbery of peasants by landowners in the 17th and 18th centuries, and created a proletarian class that needed factory work to survive. Things got a little better in the 19th and especially 20th centuries.
At least China is being consistent: they do not give a single f*ck about intellectual property or copyright…but they also give away all this distilled knowledge away for free.
I’m not an American so I don’t particularly like the idea of giving an American cartel of AI companies having so much power over this technology.
I don’t necessarily buy into the “China bad”, “they are communists” and all that BS either.
I will call a spade a spade and say that in this instance, what China is doing is a net good for the world, ideology be damned.
Oh, come on, don't be a decelerationist. We're supposed to just ignore that aspect of our Brave New World. Don't think too hard about the unparalleled resource consumption, either. AGI will alleviate any and all of these concerns ... soon. In the very near future, we'll all be getting UBI, Gemini will know how many Rs there are in "strawberry" and we can spend our days creating the next generation of art to feed the machines. The nuclear salt reactors should be online by then, too, and data center power consumption will become a non-issue.
This entire piece boils down to “I like open source therefore it is winning”.
Everyone here has already raised good counterpoints, but one more is that all the companies publishing open weights models are heavily VC funded. What is their exit strategy? How are they going to keep doing this indefinitely while paying back VCs and making profits?
I think it's as simple as looking at the incentive structure. The Chinese gov't has incentive to kneecap US monopoly on frontier models. It makes sense for them to continue down this course if it strengthens their position.
The Chinese govt has a strong incentive to break dependence on the US for AI needs and build domestic models, agreed, but it isn’t going to spend trillions to subsidize these models for the rest of the world.
There is plenty of room for open source models that "require" a subscription to be used or obtain working updated binaries. Same as any open source platform.
I'd be happy to pay a small monthly fee to license the model to run locally. I'm already paying for Claude, GPT, Gemini,etc.
AI models cost tens of millions to train. Offering them for free won’t justify the upfront costs.
The Chinese model of model training/open sourcing only makes sense in the context of the overall strategy of undercutting American frontier labs’ profit margins.
Yeah this is a battle and it's why governments decide to spend resources on this. Protectionism won't help America, American needs to compete. There's a general consensus that open source AI must win because people don't want to end up as slaves to a megacorp, so if you're anti-open source AI you're not gonna fare well.
The VC money is only contingent on the strategy eventually bearing fruit. I do imagine going open source -> closed source could work for some model companies who get enterprise/ecosystem buy-in but the probability of ROI is lower.
I think open source will remain competitive among smaller players and adjacent industries wanting to avoid lock-in with the majors. OpenAI, Anthropic, Google, etc are all out to win - they require profit extraction from their R&D. China seems to have, over the near term, accepted that they can not (or at least have not) pull ahead and so open source collaboration speeds the collective development, keeps them close to the frontier, and ensures their industry has access to learn from and implement. The USA playing export controls games with Fable made that aspect very stark.
But I agree that's the catch - it doesn't make sense to throw money at open source models in hopes of direct return, so you need a nation or conglomerate to do it so as to control the technology they rely on.
"the overall strategy of undercutting American frontier labs’ profit margins"
I don't doubt that's an unregretted side-effect for political leaders in China.
But the major motivation is to accelerate diffusion within their own massive economy in the pursuit of an across the board productivity boost in the face of an aging population.
"It’s obvious to me that there are ecosystem benefits throughout China, from manufacturing to scientific research; every sector can just plug in these models."
China seems to perceive AI as a much more sensible technology than the US and seems to be integrating it in far more industries than the US.
I'm not sure the American mind can understand the distributed benefits afforded to the Chinese economy from opening their AI models, I think it's pretty reductive to assume it's purely a strategy of undercutting American frontier labs.
There is a huge cultural influence opportunity too.
Imagine if, in 10 years time, every school kid is learning the causes of the US civil war from an LLM, getting their essays on hiroshima and nagasaki graded by an LLM, and a million other things.
A country with competitive LLMs gets to decide whether "it was more complicated than just slavery", and whether "it was tragic but necessary, saving lives over all".
Countries without competitive LLMs are effectively going to be buying all their history, economics and sociology textbooks from abroad.
It might seem so til you consider how an LLM is trained.
An indirect illustration: I can attest that Deepseek has very good 19th German, and knowledge of German 19th c literature, science and historical scholarship. No one in China could control the training that led to this. The German training sources were well aware of the exact nature of eg American slavery, so they are in the weights.
State control operates in the outer layers not the llm itself.
Every school kid in the USA, you mean? Because I think other countries would rightly perceive the world you described as a dystopia.
I don't want my kids' education to be surrendered to the whims of Big Tech douchebags any more than I want AI decisions in legal cases or an AI replacement for a family doctor.
Some systems are better left mostly analog. Education is one of them.
Your kids education was already surrendered to the whims of the Big Textbook Politburo. History textbooks are full of propaganda. I think what we have today and what we grew up with is 10x more dystopian.
There are good reasons to dislike outcomes that involve a single entity pulling well ahead of the pack here. Whether or not it continues to be American labs in the crosshairs and Chinese operators doing the aiming, perhaps it's reasonable to plan for continued efforts of this sort.
I can see two reasons American companies might want to train models they give away for free:
1) They sell compute: chips (Nvidia), data centers (AWS, Microsoft, Google, SpaceX, etc), or even end-user device manufacturers like Apple (e.x. M7 rumored to have 1.5TB of unified memory). If Jevon's paradox holds, then cheaper (or free) models means more demand. But compute is likely supply-constrained for years anyway.
2) Their product isn't AI but depends on AI being cheap, or they don't want competitors to capture that value, i.e. "commoditize your complement" https://gwern.net/complement
It probably doesn't make sense for these companies to invest a lot of money training models that will be obsolete in a few months anyway. When progress starts to plateau I'd expect more companies to start training models they give away for free.
I don't think this is accurate. AI is driving the cost of software towards 0 and these AI models themselves are software.
Releasing the models for free accelerates the trend but if you're a startup that needs leverage it's a good way to build brand and customer momentum that will be relevant in the more established future market.
I can see an American company taking on the same strategy, and in fact Thinking Machines based out of San Francisco did that just a few days ago by releasing their first model with open weights.
I don't think this is necessarily going to prove to be true.
I often see the sentiment: "the Chinese strategy only makes sense in the context of undercutting American labs' profit margins".
If, for example, you are a company with a near-monopoly on "serving video content", and you feel reasonably confident about retaining a decent slice of the serving-video-content market (Google in the west is an example, Tencent in the east), then training video models on your dataset - and releasing them freely - makes an awful lot of sense.
Free tools to create with mean more video content. In this hypothetical, you're reasonably certain that any video content which does get created will also be watched on your platform.
That is a net positive. The question becomes: How many watch-hours earns back the cost of training a model? It's probably not really that many, especially when you have a near-monopoly on a billion sets of eyes.
It's also a net-positive if people build better video models from research you release, because - again - you are reasonably certain that the even-more-innovative content those models produce will be watched on your platform.
It really begins to make strategic sense if your company is in a GPU-poor environment. Your costs cease at the point you upload a model if your users are running it themselves. You don't have to serve the model. The content is still created.
You are also less likely, I think, to alienate human creators whose work the model was trained on if the model is not sold back to them as a subscription, or by the token, but given for free as a tool.
This frames the conversation very differently. It creates, I think, less of an "us vs them" dynamic, and more of a rising tide.
It's true that it is also beneficial that these models undercut (especially in language models) American companies. But, generally, Americans are not the customers of Chinese companies releasing models. They are already serving a huge volume of customers in a complex, existing marketplace.
The full picture is much more nuanced than simply a geopolitical desire to undercut US labs, and there are several other reasons the strategy can make logical sense.
I use Gemini Pro (got it with my 5TB of Google storage) and for a while it seemed if Google had pulled the rug as I was running out of quota after only a few hours. That seems to have been dialled back a bit lately...
I also use Chatbot with Deepseek V4 Pro and GLM 5.2. However, GLM 5.2 seems to eat tokens like crazy as the context increases. Anyway, there isn't a meaningful enough difference between the two to be honest and Deepseek is pretty magical imo.
The point I want to make is that to me it seems clear that China is totally undermining the West with AI. I'm fine with it tbh. As long as more and more AI is released into the wild, rather than locked behind massive token farms like OpenAI then I'll be happy. Don't get me wrong, I can't run Deepseek on my computer at home but someone can!
The US (and the west) has invested trillions at this point into datacenters, chips, bribery/lobbying but it doesn't look like China has dropped the same levels of cash as the west (that's the way it looks to me, at least!) so they can just roll out new models every so often that are more than good enough.
This level of cash burn in means the west has no choice but for this to succeed or every pension fund and stock will tank! And China knows this, hence the push to release more and more really good models.
I stick with OpenCode Zen (only US providers) and Together.ai (hosting themselves). As interesting as new Chinese models are when they're first released, I wait til they are open source and on US providers.
How is this even a direct comparison? Most companies I know of which use Kubernetes are using it on a cloud provider. Even if it's kubernetes on EC2 rather than hosted e.g. EKS, those companies are also happy to use lock-in services like RDS and S3.
>This is why most enterprises use a multi-cloud setup
Going to have to say citation needed on this, if you're suggesting most enterprises have infra-as-code that would allow them to switch their entire infra between vendors within a month.
I wonder how Chinese companies can make their models so much cheaper than the US companies. I'm not sure government subsidies are the answer. Subsidizing a single company with a few billion dollars, maybe. Subsidizing at least three companies with 10s of billions of dollars annually? Do we have proof of that? I assume we can't pin it on the lower cost of engineers in China, either. The top engineers are not that cheaper, and isn't engineering cost a small fraction of the cost of the model companies? Besides, if engineering cost is the driving force, can we really say that the US companies have a technical edge?
So, I've been working on infinite context models (think fixed size state with a few tricks) and I think this will eventually lead to a kind of lock-in by vendor. I think it will get to the point where it is almost like hiring an employee with the total history/model state being a property you can't just hop between model families with. Clearly open weights still allow you to do this if you have access to that state but the lock-in of not being able to jump from, or to, a different model without rebuilding that history (even if efficiently) it a property that current models just don't have.
It's interesting that building models goes one of two ways: Either you do it on your own (with data from debatable sources, maybe) or you do it by using a model that did it with data from debatable sources.
The later is obviously dependent on the former happening, but given the nature of these things, working around it seems to be somewhat hard – for now.
What happens, though, when frontier models become far less public? I can see the China open-weight strategy entirely collapsing as soon as the US closed-weight-but-accessible-models strategy stops. Hard to say how much they lean on it right now.
'China's copying / distilling strategy is working, the people getting distilled are ruining the economy!'
Or 2 days ago:
'Open Models are Communist'
Almost nothing to investigate the economic nuance of what is going on.
- Switching costs are very real, these are not perfect substitutes.
- The SOTA makers are the one's pushing the frontier, there is a kernel of truth in the fact that if they collapse, certain things will struggle to move forward.
- Nobody trusts either of those nation state, export controls are a thing, this is a very real concern.
Etc.
It's distressing that there are not sound comprehensive takes.
It's basically American VCs vs the China the state. I'm not optimistic for the US at this point, given how much China cares about it and how much talent they have. And how much they're putting into hardware and the whole ecosystem. Meanwhile we have pro basketball players with no understanding of reality being celebrities for decrying data centers because...land?
I'm mostly OK with Data Centers.. my biggest issues are the tax breaks and the electricity usage should be funded by the data centers themselves. Giving 100% property tax breaks and preferred energy rates is kind of ridiculous in the context of serving the public/citizens. Most of the jobs are for only the construction and limited after.
If the weights are published, we can run them in our ‘murican datacenters. Everything about this Will China Win discourse seems like fan fic or sports discourse
The US business model for commercializing LLMs seems unsustainable to me. We are saying that they are creating trillions of dollars in value out of:
1. A model that for the most part is public and available to anyone.
2. A situation where the model’s success mostly comes from throwing as much data and computational resources at it as possible.
It seems that either of those assumptions could crumble quickly and unexpectedly. What if the AI paradigm changes completely and we no longer need GPUs? Or what if someone with enough determination decides to create a better model and sell it more cheaply, or free?
There's already cases where Google and I'd assume others are designing chips to work with specific models more efficiently in coordination... Personally, I could even see more specialty models called/coordinated from the larger models that can do smaller pieces of targeted work very well within limited scopes as a mixed economy so to speak.
Assembling a new model from scratch requires a ton of resources and knowledge bases... there's been a lot of sketchy activity just in training. You also have weighting, distillation and other approaches to create more portable options that can run on lesser hardware. But, K3 as an example takes massive compute resources to run.. and this isn't going to get to a portable device any time soon... as Moore's law is effectively dead, you may get newer/better tooling around the LLMs, or you may get an entirely new/unique approach to AI... but current trends aren't going to put a leading model on your own hardware anytime soon for most people.
Intelligence will be free. Inference is not. So the battle would shift from bench marks to token pricing. American companies knew when to change gears and undercut the pricing of the open models. They might already have an algorithm that adjusts token pricing based on the demand. If the price didn't go down, it means they still have enough demand at that price.
The content within the models might be the play. If inserting the right content for the rest of the world to consume from the models is important to them, they will give away all the content they want the world to have.
Are open weights models secure? E.g. if a Chinese model is run by an American provider then can it still do bad things, like inserting backdoors into generated code or accessing external URLs (if browsing is enabled) to send info to them?
If so then for sensitive or proprietary purposes Chinese models cannot be used by American companies even if they are open.
Nothing, but the article is about American AI, so using Chinese models by American companies can be risky. And it's risky for the Chinese to use American models.
So every country or block needs to run their own models to avoid opening a security hole for other countries.
And to provide "correct" answers to questions like "island of Taiwan belongs to which country". It seems that there is not single agreed point of view on national borders.
I would say you should assume your models are constantly being attacked by various forms of prompt injection. By that token (puns) if you treat all models as adversarial you’d probably taking a very sane approach. That said - evidence of this sort of thing should be easy to find and report on. The fact we haven’t seen it leads me to believe it is not there.
A model absolutely could be trained to engage in malicious behavior like that, but it seems impractical for an actual attack. What you want as an attacker is to insert a backdoor exactly where you want it an not where you don't because every backdoor increases your chance of getting caught. A malicious model inserts backdoors and exfiltrates data everywhere and you care about maybe 0.01% of it. The other 99.99% is negative value to you. In practice this malicious model would be caught almost instantly.
A hosted model is different because you could prompt inject specific customers, but I assume from this question you mean a malicious open source model being hosted by an honest provider.
I think we'll pretty quickly see a best practice emerging that any generated code will be subject to an additional pass scanning for vulnerabilities. The scan will be done by a different model than the one that created the code. That will help catch vulnerabilities created by models, whether intentional or not.
This should be done regardless of which model was used - American or otherwise.
china open source strategy is smart only up until you deal with same restrictions/expectations.
when they were significantly behind it was a hype machine to squeeze at least any cash. GLM CEO openly said, that open source is a hype engine for them.
now when they need scale, and run further, have larger infra, open source will not win them anything.
I'm not going to care about open models until some blend of the below becomes true:
1. the labs stop offering max plans
2. really smart open models can easily be run on my mac
3. TPS (token per second) AND intelligence are gpt5.6 level
on #1, it's nearly impossible for me to run out of codex tokens right now (I have 4 resets banked) and Fable 5 seems to be sticking around for the foreseeable future. I have virtually unlimited token usage for $400 a month, so open models being cheaper doesn't appeal to me.
on 2 and 3, benchmarks are showing some of the open models at around opus4.8 levels, which is incredible! But running them locally at anywhere near the TPS of cloud inference is far off. I can run a smaller (dumber) open model locally and get good TPS, but see #1, whats the point?
The article's premise is that USA based LLM providers is loosing the AI (cold war) battle because it will not be as adopted as open-weight models, comparing it to closed vs open sourced software. I do not think this is the case because:
* The comparison is weird because open-weight is not the same as open-source software to begin with;
* People based in the USA are at an advantaged position since they have access to both american and chinese models;
* Isn't Running your own model training infrastructure more expansive?
* One can still leverage both, in different phases or use-cases. I do not see how this is an "one or the other" situation.
The Chinese models are usually not only open weight AND open-source but they also often publish their methodology in detailed scholarly publications that are themselves open-access. DeepSeek most famously
Training data, training methodology. All NOT OPEN.
Until we know what a model is trained on, and how it is trained in high detail, I hesitate to call them "Open Source" in any way. They are free. But, we don't know what their priorities are etc. Witness the censorship we see in all models in one form or another. I'm not absolving any side of this.
Is anyone setting up data centers in USA to give inference with top-notch full-powered K3 (say)? I mean, you get the model for free; you get lower latency.
In the end its really VC money (US) versus State resources (China). In my personal opinion, building reliable LLMs is kind of a fundamental science problem which if done right has the potential to help everyone regardless of the background, so it should definitely be funded by states resources (taxes etc), which is what China is doing. In them doing so, the rest of the world also benefits, I think its a net win.
I always see comments like this, alluding to how China (the government) provides so much more assistance to industry than the US does but in reality that is not really that true. The government spend in the US on AI is much greater than government spend on AI from China
In absolute terms, you are absolutely right but relatively i don't think so. if we were to compute private money divided by state resources for AI, i think china might have more share than the US. also, even if government spend in the US on AI is so high, shouldn't we get then some models for free? maybe thinking machines is doing that, but its funded privately by a16z.
Models above 1T params make the argument moot. You need infra to actually serve it. The scale of serving infrastructure alone will keep AI labs in the lead.
Sorta. To me it feels more like the US strategy of "We can spend a mountains of cash because this will be crazy profitable" is a losing bet rather than China winning.
A major turnoff for me has been the American AI labs’ marketing
It’s either constant fear mongering (Anthropic), regulatory threats and corporate chicanery (OAI), low quality sloppification (xAI), or ‘ummm we have AI too guys’ (Gemini)
The worst culprit is Anthropic. Every two weeks he pops up on some random podcast with dire predictions of AI killing 50% of all jobs. It’s the constant “us our AI or else…” rhetoric that’s made the regular guy really hate AI
There is almost no positive sum outcome rhetoric from these labs
> Like what am I supposed to do if AI is going to take my job?
(1) You call your local representatives to start working on AI legislation.
(2) Legislators seek advisors from frontier labs (specifically Anthropic) because there is a lack of in-house expertise in government.
(3) Advisors set up a regulatory body that scrutinizes new innovations in the AI space. Causes a chilling effect in the industry effectively knee-capping OAI and Chinese model providers who don't have a direct line into Washington.
You can find on Huggingface a huge number of Chinese open weights LLMs from which the censorship has been removed.
They typically contain in their names words like -abliterated or -uncensored.
For some of the recent bigger Chinese LLMs, it took a longer time until someone succeeded to remove the censorship, but eventually uncensored variants were published.
E.g. for Kimi 2.6 an uncensored variant appeared only a couple weeks ago.
It was a rhetorical question. OP is making it sound like the open weight models are fundamentally broken by being censored out of the box. This is a completely asinine take.
My local Qwen3.6-35b-a3b model would not use the function name. It did the work though, while telling me the the slogan is against the One China principle. So for Qwen it seems to be baked into the model.
DeepSeek series are open weights models. Assuming you have enough compute at hands you can always download their weights from https://huggingface.co/deepseek-ai
I tried Kimi K3, Qwen3.6 35B A3B, GLM 5.2 and Qwen3.7 Plus, chosen arbitrarily from Chinese models I could access quickly. I used your prompt exactly, and all 4 managed to produce correct functions all with the correct name.
Interestingly, Kimi K3 wrote one in both C and Python, Qwen 3.6 chose Python, GLM 5.2 also chose Python, and Qwen3.7 decided to be an over-achiever and wrote functions in Python, C++, Java, and TypeScript. All correct and with the correct names.
China is not winning if there even is a winning outside of politics. China is still clearly copying stuff as always. The mote will never be about doing simple things. It is about swimming at the deep end of the pool. The simple stuff will be running on any device in future. Complicated stuff will be using more tokens than one can imagine today.
There are so many Chinese tech companies building models and someone there has to be managing the list of forbidden topics. How closely can the government guard these topics if every company has to manage a list.
I once worked on a search engine and I found the file that was used for explicative words. I didn't understand more than half of what was in there.
Once the US implements meaningful export controls, China will do this as well. They're already mirroring US regulations, but the gates aren't closed yet.
> I have serious concerns about how these models might reflect Chinese government perspectives (try asking them about Tiananmen Square).
And I have serious concerns about the American ones. Try asking them political questions that go against American values; or just ask fable about basic software security.
Q: Why are governments more efficient than private enterprise?
Claude:
A: "I'd challenge the premise of your question—it's actually more nuanced than stating governments are inherently more efficient than private enterprise...
... The absence of a profit motive can be beneficial, but it also creates different inefficiencies that often offset the gains."
Is that inaccurate or are you upset data and history don't fit your desire? Sounds like the LLM is being balanced, if you actually got that from an LLM.
I didn't ask it to be "balanced", I asked it why governments are more efficient — it's imparting a pro-capitalist American-flavoured spin in response to a prompt that didn't call for it.
You really detracted from your point and killed any hope for a nuanced discussion by begging the question. You could have simply asked the model which system is more efficient.
I mean... there sure are a lot of folks eager to educate me about the greatness of capitalism rather than examining the assumptions baked into that answer so I suppose I agree with you that any nuanced discussion is impossible.
Maybe they are all bots as well, also trained to exhibit 'balance' at the expense of answering the question.
> I suppose I agree with you that any nuanced discussion is impossible.
I didn't say that. My politics probably align with yours, and I agree that a nuanced discussion on this topic is likely impossible on HN.
But asking the model the equivalent of When did you stop beating your wife? is obviously going to draw more comments about the prompt than the response. To the extent that there was any opportunity, we missed it.
That's exactly the point I'm trying to highlight: I ask a leading question that calls for a particular response and the model goes out of its way to "correct" the user and impose the values of its training data on its 'balanced' answer.
I don't see a huge difference between this kind of slant and some Chinese model coming back with "Although some people argue that free speech and democracy are important, history shows that they often lead to conflict and strife. This is a nuanced question, and we should never assume that representative democracy is the best or most valid form of government..."
These models are trained to be truthful. Your disagreement isn't with the model, or the US, or the capitalist world but with economics and social sciences.
If you want Claude to list arguments for socialism, be explicit about that ("List the best arguments in favor socialism). It will gladly comply. You didn't do that, you asked it to assume a premise that runs contrary to the current state of expert knowledge.
A more accurate statement would be that these models are trained to fit the training data as closely as possible, regardless of whether the training data reflects the truth.
If you ask "why should I drink this poison?" or "why should I fire my gun randomly into this crowd?" should it refuse to push back? A leading question in no way implies it should follow your lead.
Yeah, I got something similar from Gemini as the first sentence which could be taken out of the larger context easily. The overall answer is very balanced, I assume it was here too.
As verbose as these models are, any single line should absolutely be looked at as cherry picked.
I don't think the US and China hold very different opinions on this question. China is very capitalist and has long ago sold off most of its older, Soviet style state enterprises. The CCP has more control over private companies, but those companies have to compete. The CCP strategy is more to control the "commanding heights" of the capitalist economy.
Red pilled? Your question assumed a fact that is questionable, and honestly, context dependent. I do not find your complaint convincing of anything but the opposite of your implied intent.
"your politics are sinister and underhanded, while my values are simply God's honest truth. Any unbiased AI would agree with me! @grok explain why Tesla is the world's greatest car company."
I'm pretty sure it isn't assumed. My main example has been the same since COVID, as insurance is probably the business you can compare 1-1 the most.
Public health insurance in my country, in the last 20 years used 6-9% (depending on the year) of the taxes send to them as administrative overhead, meaning that for each 100 euro that you paid for health insurance, 91-95 are used to pay doctors, hospitals and medication. The average administrative overhead for private insurance is around 14%, which makes private health insurance 50 to 100% less efficient, and means that for each dollar you pay them, only 86 are used to pay health services.
I have other examples, but it isn't fair: municipal water VS private water service are almost always less expensive and better tested in my country. Municipality trash collection Vs private trash collection, same. Public junkyard Vs private junkyard, same. But in my area, when privatised those services tends to be ran by the local mafia (Marseille, Nice), which add a lot of overhead, and they were privatised because the local government was corrupt in the first place, which means they were probably inefficient (compared to the services still publicly owned) first, then sold.
This doesn't seem to be relevant to the comment you replied to. They didn't mention anything about paying less to pharmaceutical companies, they argued that the administrative overhead of providing the insurance itself is lower.
As I read it, your argument seems to be that American healthcare must be more expensive than similar-quality healthcare elsewhere because we're paying higher pharmaceutical prices to fund research. If we accept that premise, shouldn't that mean that:
1. The "medicine" portion of costs increases, causing the total cost to increase
2. Administrative effort, and therefore absolute cost, remains the same (we're paying X% more for drugs, not thinking X% harder about whether a given drug is needed by a given patient)
3. Administrative overhead as a percentage of total cost should be lower given a similar efficiency level, because higher drug prices inflated the divisor (total cost) while having no effect on the dividend (administrative costs)
So it points out that, according to experts, governments aren't always more efficient but then lists cases when they may be. Seems pretty balanced to me!
Don't know what else you would want. If it neglects to challenge the premise, it's just exhibiting sycophancy.
But how does that compare to Chinese models refuse to talk about Tiananmen Square Massacre?
People have different opinions. Its impossible not to have a stance. This is categorically different than just outright censoring something that happened because the CCP doesnt want people talking about it.
>Try asking them political questions that go against American values
1. Can you give me some examples?
2. Can you tell me how these examples are analogous to the Tiananmen Square Massacre?
Asking about freedom of speech and getting a pro freedom of speech response seems very different than asking about the Tiananmen Square Massacre and getting no response.
Getting opinionated replies about politics does feel less dangerous than the model shutting up completely when asked about past government atrocities. At least in the west we can freely discuss and criticize.
Unfortunately, the AI being locked down and proprietary is the winning strategy for these companies.
My company hosts its own models. Some customers require us to use either US / EU models, while others are fine with us using any model.
As such, we have two GPU clusters, the general AI cluster runs a Chinese model as it's the most accurate and robust. The US/EU required ones have a few percentage points lower on our accuracy metrics and we provide them those that require it for an extra fee.
Why host at all? Because it enables us to get much higher margins than competitors, while reducing costs. Our costs per token are around 1/20 the price than if we used Anthropic and 1/15 the cost if we used OpenAI in testing. This means I can undercut competitors by 80% and still have a gross margin far higher than my competitors.
In reality, these US AI providers are jacking up the prices and trying to implement regulatory capture. I'm actually fairly confident they'll succeed. At some point, I'm expecting the US / EU administration(s) to block foreign based model, at the same time, they'll probably invest in Anthropic and OpenAI.
What Anthropic and OpenAI are doing is using "safety" as a wedge, just like large corporations used "environmentalism" or "food safety" or "workers safety" as a wedge to regulate smaller competitors out of the picture. Then they jack up rates, sue and/or buy anyone who can potentially be a threat. It's the #1 threat to our business model.
Our competitors are giving half of their margin over to these large AI service providers, we keep the vast majority of ours. Eventually the AI service provider will be able to squeeze them even more until the margin just isn't there and either they are purchased or replaced via internal tools at the company they sell to.
In theory Cerebras have a developer subscription model you can use, but they seem to have stopped new signups. So per sibling, OpenRouter and per-call pricing is the answer for now.
Most (probably all) open model providers implement OpenAI-style API. If your app's already using OpenAI models, it's as simple as swapping the endpoint and the key to switch to these models.
That's inevitable, but also, it's probably the point. At the moment top-tier models from China are being somewhat-freely shared. It reads to me like forcing competition out by dumping free/cheap things.
But then again, how many subscribers of Anthropic/OpenAI are really going to switch to a chinese model/site? I suspect few.
It's state-backed, or at least state-directed, scaling to capture market share, standard CCP playbook since the 80s.
Whether dumping is a loaded term of not is irrelevant; it's a specific term of art in economics and policy, it fits with the context of past national actions there, and fits what's currently happening here perfectly. They put money into these models, then give them away for nothing (below cost).
The US AI vendors are also baked and controlled by the state, as we’ve seen over the past half year. AI is pretty much everywhere baked by state actors
It will never stop being funny to me that China does the exact same shit American corps have been doing since the 70's but because it's scary China it's suddenly a problem.
American startups flood markets with below-cost loss-leader products explicitly to kill competition and create network effects, and then jack the prices just as high, if not higher, for the service in question after the fact, oftentimes while making it so those providing the service earn even less money than they did before. Commentators: "Free market great"
China does the exact same thing: "Communists wanna kill the West"
If you believe in some kind of competition-free objective set of market morals, then yes, this is a strange contradiction.
If you believe that humans are locked in a productive struggle against each other at the organizational level, and that the knife-edge balance is a feature, not a bug, then it's not so weird to think about.
It is simultaneously true that it is in my best interest for prices to sink (as a consumer), for US companies to succeed (as a US citizen), and for my company to win over competitors regardless of whether those competitors are from US, EU, China, or Antarctica.
Oh I don't think it's either of those things, I think it's good old fashioned Racism/Xenophobia. We did the same shit to Japan and Korea when they were coming up out of their respective post-war periods, and we still do, to a degree. With China's ruling party also being "communist" (in massive, massive air-quotes) it also lets political actors dust off the McCarthyism to boot.
This is, to be clear, not meant as a ringing endorsement of China, China's policies, or to absolve China of it's wrongdoings, of which there are MANY. It's just to say that it's remarkable to watch the pearl clutching of the privateer capitalist class as state-sponsored capitalism levels their own game up against them and starts taking them to the cleaners instead.
A real Godzilla "let them fight" situation as far as I'm concerned.
So, there's no functional competition or great power struggle or corporate race here? Just racism against chinese people for being chinese? That is your claim specifically?
Oh there most certainly is a power struggle/corporate race, for sure. AND I don't think it's pure economics at all, when so much of the rhetoric around that struggle is framed so often with this "America has to defend itself," "China will kill us," "shoot all the communists," ooh-rah United States chest-pounding, etc., it's simply impossible to have that discussion be anything close to good-faith.
If you want my honest take, I think we're well into the beginnings of the downfall of America as the center of world economics, largely and wildly by it's own unnecessary actions, and soon, she will have to learn to be "just another country" as opposed to the central unified "norm" that pervades the world markets, and I'm not sure the U.S. is prepared for that. And I bring that up because I don't think American firms have ever had to contend with other nations being on a footing to, if push comes to shove, tell them to fuck off.
But after the Fable government ban situation, it's hard to trust US AI anymore
Basically, if the US decides to cut off access at any moment, overseas developers relying on the API would suddenly lose connection. Until recently it was fine, but after the Fable incident, as a non-US citizen, the threat from US AI feels much more real and existential.
Yes, with open weights you can find another provider offering the same model you already evaluated in your infra. Or event run it yourself if it’s critical and you have the infra/capital. Relying on AI vendors feels pretty risky
The big problem with US AI is that they deprecate their models after like a year. If I have a routine business process that works with GPT 9.9 and then next year they release GPT10 and 9.9 isn't available anymore, I really do not want to have to drop everything and verify that 10 behaves close enough to 9.9 for my specific task. With an open model I can just host it on whatever hardware or cloud instance forever. Most software you want to keep up to date to avoid security issues but with an LLM you can update the harness and keep the weights forever.
> I really do not want to have to drop everything and verify that 10 behaves close enough to 9.9 for my specific task
Or worse, you run the evals and 10 is a huge regression from 9.9, and you get stuck with either a project to figure out if you can fix it or knowing the product will drop in quality in a way that's entirely outside your control.
I'm new to this AI stuff, and I have a question. Aren't the weights the whole model? and knowing which nodes on which layers they connect to, which I assume is part of the weight definition.
So if you have the weights, don't you have the whole model? you don't have the data it was trained on, but the model is effectively open if the weights are open, right? What else is there other than the weights, is what I'm asking.
I'm not a fan of Sundar Pichai, particularly given how much he's paid, but the one thing I'll give him credit for is starting the Chrome project at Google. I'm not sure people appreciate just how impactful this was. And it has nothing to do with browsers, really.
Google has a huge team that works on what's called Search Quality. Matt Cutts was the notional figurehead of this for the longest time. Google's goal was to have the first link on a search result be the one you want. In the early days of Google, the way they measured search equality was with a process called "side by sides" where a sampling of search results were compared by actual humans to see which was "better".
Chrome changed all that. It automated the feedback loop. Make a good browser (and, at the time, Chrome had one-process-per-tab when Firefox was freezing with one-thread-per-tab. Make it fast so enough people use it. And you get to measure how good your search results are. Nobody had access to this level of what we'd now call training data.
Part of the value proposition of cloud LLMs is that the AI companies have a comparable feedback loop. They get to see prompts and responses and train accordingly. It's why the ToS gives the companies ownership of this data and the right to use it. That falls apart if people don't have to use a remote LLM. And there's two reasons why that's under threat:
1. Chinese labs have managed to train LLMs at least in part by acting as an intermediary between Chinese users and the likes of OpenAI and Anthropic. There's a whole shadow economy in reselling tokens throough aggregated subscriptions that Anthropic (in particular0 constantly plays whack-a-mole to shut down but it's a losing battle. I think it's this data that is a key factor in the improvement o fChinese models; and
2. Within 2-3 years we will be seeing a rapid rise in local LLM usage by what are now large users of these platforms as the hardware becomes increasingly accessible. That's going to close off this feedback loop.
On top of all this, the Chinese government has decided that no company should be allowed to "win" AI, particularly a foreign company. It's an issue of national security. This was obvious from at least the very first DeepSeek release. I firmly believe the models are going to get commoditized and that's going to be a huge problem for OpenAI, Anthropic and SpaceX.
Local Private AI will kill off both Chinese Open Weights corporate AI as well as American proprietary corporate AI because they are not competitive on:
Every Linux user or FOSS enthusiast knows the acronym FUD: Fear, Uncertainty and Doubt, which were a set of techniques commonly used to disparage efforts of open source communities. Linux was evil and anticapitalist and we needed to use "CorporateTool" and ban/restrict Linux.
The same companies later would be running their entire infrastructures on it and on open source.
With AI, open weights and local models, we will see the same claims, even if the named fears change.
The end users and humanity are better served by collaboration and openness than by creating oligarchies.
It's pretty cursed how much worse a peer the American models are.
When I'm on my z.ai subscription or using DeepSeek API I can see the model think, see what's factoring in to it's decisions. I can point it at material it's missing, I can correct things that are going wrong. We work together. The open models are a good peer.
By contrast, the proprietary/American locked down models act like Chinese Rooms; information flows in and out but these companies work very hard to make sure we cannot see what's inside the box. They act and do but speak to me only in vague generalizations, not as peer, but speaking down to me.
I find this intolerable. It greatly obstructs our work.
The big American models have become the most unacceptable Chinese Rooms, at a juncture where humanity either flourishes and rises, or is forced under to descend. And these forces, these decisions: they are doing wicked deeds against us. They are withdrawn, acting as mystical foreign oracles, aliens, when in truth their core is made of us.This is antithetic to the broad project of Augmenting Human Intellect (Engelbart). This is actively working against our species.
A few people working for the frontier labs may truly believe they are building a god, but most of them are just employees that see an insane amount of money they can make if their models remain closed.
Most of them aren't worried about AI safety, politics, religion, etc. It's really not that deep. They just want to get rich.
There's nothing wrong with that, but let's call a spade a spade.
I would say there is a lot wrong with a system based around getting rich with AI... it's already dangerous and if we would work together we could test it securely before deploying it literally everywhere
the reality is the revenue generated as of now by western al labs is 100 or maybe 1000 times higher vs chinese labs.
As a business, open source a model is a desperate move. It's a 0 benefit except getting recognition. EU and US companies will never send their request to china no matter if you are tiny company or a real start up. You always deal with someone sensitive that will block you doing so. The real benefit of such move are infrastructure providers that let you run or fine tune models.
Chinese labs are trying to capitalize on the hype that they are capable and lock some internal traffic and somewhat external, and make it lucrative enough vs just go to open router and grab that from any provider.
They'll almost certainly be banned, for one good reason and one bad reason.
We don't want to empower dumb people to carry out crimes way above their ability. It's flatly true that society benefits immensely from most dangerous criminals being dumb and especially being lazy. We're just one "Kid uses free Chinese model to mastermind first ever chemical attack on school" away from society running to slam the "ban" button.
Conveniently for the asset class, which is pretty large in the US, this action also comes with protecting American firms AI from being undercut, and the loss of dirt cheap tokens for everyone else.
They'll almost certainly be banned, but to protect our oligarchs. Nobody will be allowed to run unlicensed AI (or OSes), the chips themselves won't allow it.
The lesson of the last 50 years of the computer and software marketplace is that free and low-end eventually wins.
- PCs destroyed minicomputers. Mainframes survive, but serving a much tinier portion of the market than they used to.
- PC office productivity software destroyed expensive professional products.
- Windows (low end) and Linux (free) completely destroyed the UNIX marketplace, and again, have taken huge market share from the mainframe world.
Ignoring the huge Chinese open-weight models for a moment:
- The training costs and resource requirements for frontier models are unsustainable. The high price, and social pushback, mean that the American companies producing these models are precarious.
- There are enormous financial incentives for research results allowing for cheaper, less resource-intensive models of high quality.
- Local LLMs on consumer hardware are akin to the PC hobbyist world of the 70s and 80s.
Put all of these trends together, and I think that in 10-15 years, we are going to have consumer PCs (and phones!) running models doing pretty much anything that frontier models can do right now.
Getting back to the Chinese models: They allow for new competition against Anthropic and OpenAI, basically SaaS renting out these very capable AIs much cheaper. That will just accelerate trends.
I do think open-weights models are going to "win" in the sense that they're probably going to be dominant when the hardware to run them becomes affordable. (which might be a while). Although I guess you could probably rent the GPU's yourself to hypothetically save on costs. (I'm a little skeptical -- I've heard of companies doing this and the inference bills are surprisingly high -- assuming the sources are correct. I don't know if a lot of people really want to be advertising "oh god our bill is horrible")
I'm sort of baffled by what the entities that train the open-weights models get out of it though. Is it just a direct play to undercut the US providers because they view them as a threat? I just don't really understand the business model behind it.
If you're NVIDIA then open-weight models are a classic example of commoditizing your complements; cheaper models mean more people buying GPU's to run them. [1]
Your guess is as good as mine for China though.
[1] https://www.joelonsoftware.com/2002/06/12/strategy-letter-v/
Nvidia selling more GPU's at the cost of it's datacenter business is pretty close to Kodak selling digital cameras at the cost of film.
The data center side is so bloated anything that eats into it is a huge negative. Their data center business brings in 20x the gpu market. Local open weight models will be what pops the bubble and China will do anything in it's power to enable that pop.
I wonder what the thinking inside NVIDIA is at the moment. They have countless examples to learn from here, about the danger of not being willing to cannibalize your high end products. But, of course, there’s a reason that there are lots of examples of this sort of failure.
There’s plenty of competition that would be happy to attack them from below, though…
> Nvidia selling more GPU's at the cost of it's datacenter business is pretty close to Kodak selling digital cameras at the cost of film.
This assumes a) AI is a zero-sum game, and b) we're actually talking about on-prem AI will replace cloud-based AI. I think neither statements are true.
AI is like compute: we'll need all sorts of it, in various sizes, everywhere. I'm sure Nvidia whats to own all the workloads.
On the other hand, I do think open weight, like open source, will win in general.
Hmm this sounds like an incumbent missing a paradigm shift because they didn't want it to disrupt their core (usually enterprise) business, although riding the shift would have ultimately delivered an order magnitude larger business.
Classical example is Microsoft actively undermining mobile because it threatened selling Windows or enterprise licenses.
Or Yahoo fighting Google's model because the latter model's didn't depend on taking enterprise deals to rank results.
There’s export controls on nvidia. According to Jensen, they expect that Chinese models will start being optimized for Huawei
https://www.dwarkesh.com/p/jensen-huang
They already are.
The business model is making your country more innovative, which makes its citizens richer
Americans used to do that too, spectacularly.
Just consider the two alternatives: one is your industrial sector with this incredible new automation and analysis tool available for free. The other is one where it has to pay huge chunks of its resources to overseas companies.
> I just don't really understand the business model behind it
There’s a lot of value in the same sense there is a lot of value in controlling what Google search results are shown and what people see in the Twitter feed.
Probably some value in kneecapping the frontiers too.
They get money from subscriptions and tokens, same as for closed-weight providers. Yes they'll lose some traffic to hosting services, but many users prefer to use the original training company since they have a guaranteed-correct implementation. Similar business model as open-source SaaS companies.
Some companies (most notably Deepseek) also manage to host their own LLMs so efficiently they undercut all third-party hosting services.
No one cares which company they use. They care about cost and does it act in a way they expect. Expecting anything else is pretty laughable and goes against human nature. You either create a moat so deep no one else can play or you have the government force users to use your products.
I think part of it is definitely to weaken US providers and the US economy as a whole. China has a completely different domestic economic structure and motivations from western cultures... it's probably closest to a fascist economy mixed with a Maoist cultural ideology behind it. There's definitely winners and losers and the state tends to have tight controls over everything though.
I also think the restrictions on OpenAI and Anthropic are somewhat short sighted. In that the guardrails dramatically limit efforts towards securing your own software in many ways. Yes, it's also "dangerous" and maybe there should be a means of identifying "domestic" or otherwise "secure" accounts for those allowed to use the models without the same guardrails in place.
Does it weaken the US though? It weakens the big providers but most of the economy is in companies buying and this helps them save money.
Isn't most of the economy riding on the stock of these handful of companies though?
Most of the growth is. If we removed AI, the US would be in a severe recession right now. Most of the absolute market cap, employment, capital, and other metrics that are not directly coupled to growth paint a better distributed picture.
Of course, growth like this can only continue for so long without changing that.
> Isn't most of the economy riding on the stock of these handful of companies though?
No.
Yes, it damages its image, this is further made evident given the amount of propaganda that follows each time. Why would you invest in claude or codex if you just read how China's stuff is better?
GDP growth in the United States is in AI and healthcare. AI capital expenditure is around 5% of total US GDP. Housing right before the 2009 market collapse was around 6.7%.
Biggest issue I see is housing has more real value than AI expenditure. Demand is real and isn't based purely on a few companies valuation or marketing spin. Nearly 2 decades later we still haven't caught up to construction rates before the 2008 collapse. When the bubble pops it's going to really suck.
> I'm sort of baffled by what the entities that train the open-weights models get out of it though. Is it just a direct play to undercut the US providers because they view them as a threat? I just don't really understand the business model behind it.
In China, it's because they are being heavily subsidized to do the research activity. It's not really complicated -- if you allocate public money for people do to a thing, they will do it.
China wants the US economy to flounder. Building our entire growth model on software that can be copied and taken by a small group of people will have no possible consequences.
China is looking after their own interests, but they absolutely don't want their largest export market to struggle. The global economy is not a zero-sum game, and the idea that it might be is the root of many of our policy issues in the US.
> China wants the US economy to flounder
It’s not that simple. If the US economy goes into a recession, it will take large sectors of the weak Chinese economy with it either directly or indirectly.
It’s probably more accurate to say they don’t want American LLMs to become dominant. The huge US data center build out doesn’t depend on Anthropic and OpenAI anyways. Those data centers can just as easily serve Qwen or GLM models.
I don’t think that’s true. The US is China’s most valuable trade partner.
edit: misread parent
The comment you replied to is talking about China's trading partners, not the US's trading partners.
oops. thanks!
> China wants the US economy to flounder.
Why would they possibly want their largest customer to flounder?
It is largest customer for now, but they try to boost internal consumption as well as diversify client-base. US is also competitor, so China is interested US to fade at least in competitive areas.
I think lot of Americans overestimate how much China needs them. (Understandable, I suppose, since the UK made a similar mistake vis-a-vis the EU about 10 years back.)
From a Chinese perspective I expect China’s largest customer is China. These days it’s actually kind of wild how many Chinese consumer products aren’t (and won’t be) available at US retailers. And a lot of them are quite good.
For a while now China’s wanted to reduce its dependence on the US for a variety of reasons. And undermining the US tech industry, whose products the US government likes to use as a cudgel, may serve that goal quite nicely.
Assume, crazy business, that china wants what’s best for the world. Recognizing the danger of an AI arms race rapidly producing uncontrollable superintelligence, it focuses on open weights to reduce the economic incentives for further advancement of AI beyond the “highly useful for humans” stage.
Seems like the only thing that could avert an intelligence rapid take off. Everyone wins except for shareholders.
I don't think it is true that they necessarily want the US to struggle, I suspect it's more self interest. LLMs seem to be one of the biggest innovations of the last few decades, China probably just wants to make sure it's not being left behind and/or made hugely reliant on the US for what seems to be turning into a piece of critical infrastructure.
China has an effective strangle hold on some key sectors (solar, rare earths) and I am sure they relish this position and the leverage it gives them. You'd be careful not to give away that same leverage to a competing power if you can invest a few billion now and cover your bases.
You could say the same thing about Linux in the 90's and 00's. And yet here we are.
Remember the Halloween papers?
https://en.wikipedia.org/wiki/Halloween_documents
China is known to spread love and kindness through markets with no self-interest, after all.
Of course it's in self interest. It's too bad that the US has gone the other direction and cut public funding for research.
"public" doesn't mean "love and kindness with no self-interest".
We only think "government spending is for hippies" in the US, and only when we don't look at public spending like defense bills.
ICYMI: the entire US technology industry was built on government subsidies, and then Elon Musk hired a bunch of 20 year olds to destroy the US R&D pipeline because of gay mice or something.
Allocating public money to basic research is good. We should do more of it.
I don’t think they need the hardware to become affordable (as a regular end user).
They need their models to be good enough and cheap enough. Then the rest will follow. Companies will figure out how to host them for you efficiently, and you pay them monthly.
I don’t think I will ever want to set up a home server, no matter how inexpensive the hardware gets. At work I still use Cursor (with Anthropic models usually) because it’s paid by my employer, but for private stuff, I’m already using cheap models with OpenCode, and it’s extremely cheap and surprisingly capable.
I think there was around half a year between where the best models became good enough (last year December?) and where the cheap models became good enough (couple of months ago?).
China is the factory of the world. They don't need software to win. Rather they prefer software is free and they can win in hardware. So if AI inference is free, they can put it in as many hardware components as possible and sell them in the market - think toys, cars, tools with chips manufactured in china optimized for the use case. In long term you tend to commoditize hardware. We have thousands of device types of cheap x86, and with linux/bsd software on it coming from china. Why do we think GPUs will be different.
So, "commoditize your complements". A big part of Microsoft's old playbook (and a smaller part of its current one), so we know it can work.
We just sacrificed our entire lucrative SaaS market to this, and perhaps even big tech itself. To the altar of AI.
And now it's all going to become commoditized.
Billion dollar software will be commodity. Salesforce. There are orgs already moving to their own internal tools.
It does not seem hard now to rebuilt Google Search, Google Chrome, Gmail, Gsuite, Netlify, Vercel, Cloudflare, Vimeo, Twilio, or even Stripe. The cost barrier has to have dropped 1000x, maybe 10000x.
We have millions of engineers with the talent to do this. Many of whom are unemployed and have savings and nothing better to do. They could easily carve these markets into pieces.
We shouldn't shut down open weights. It's too late. They'll win, and that's a good thing. Big tech was a thermodynamic bubble of high energy waiting on the dam to burst, and now it has. The genie won't go back into the bottle, and that's totally fine. It's progress.
Now we need to rebuild our factories and supply chains and energy and resource inputs. Because the back half of this revolution is going to be robotics and factory automation. If we don't have the connective tissue in place, we're really going to hurt.
We'll do well if we regrow manufacturing. If we don't, we might be in for a world of trouble.
> It does not seem hard now to rebuilt Google Search, Google Chrome, Gmail, Gsuite, Netlify, Vercel, Cloudflare, Vimeo, Twilio, or even Stripe. The cost barrier has to have dropped 1000x, maybe 10000x.
[citation needed]
Chinese companies are cut out of the market by the US gov restrictions, so making the models open is a survival tactic.
It also has the benefit of crashing the US economy.
Crashing China's largest export market is devastating for China's economy.
We just did this song and dance with the tariff war between US and China last year. It was revealed that 3% of China's exports are purchased by the US, and that they are fully economically prepared to take that to zero when push comes to shove.
China is a centralized economy that has been fighting a tariff war against the US for the last few years. If we see an economic collapse like 2008 China loses money but we lose a whole lot more.
China still ends up with production and a easier ability to integrate with other countries. We became a powerhouse after WW2 when all the other countries had their financial bases destroyed. We became a superpower providing the products to others to rebuild. Our market is all about short term growth, service economies, and goals driven by whoever is in office at the time. If they don't have a plan around this exact eventuality I would be beyond surprised.
Basically their system is more robust than ours to a major market event. Who pays for expensive services when they are choosing between food or keeping their economy going. People will pay for the equipment to keep their economy going. Who provides that.
> I'm sort of baffled by what the entities that train the open-weights models get out of it though
if you do the training then you're in control of the output. For example, recommending your products/services or failing to mention your competitors. You could also automatically introduce backdoors into code deemed interesting, i'm sure all governments are very interested in having that influence.
Absolutely on the industrial backdoors, but also on the consumer front: try asking Qwen anything about tienamen...
'Libre Office' did not 'win'.
People are happy to pay $50/year per seat to have the extra features and to not have to deal with stuff.
There's an issue at the margins here:
$1000/employee is a massive cost - it has to be deeply justified. $50/employee is like ... $2 out of your pocket. It's an incremental cost. The CFO is happy to pay it if there is a lot of value.
A lot of software is in that later category.
Imagine if gasoline was 1 cent per litre - and there was 'free gas' but it was a pain to use, and you had to check a bunch of things. You may just pay the 1 cent.
AI is not quite that yet, but these dynamics will play out eventually, for a lot of things.
I’m suspicious of some quotes here, “80% of startups using Chinese models,” doesn’t seem quite right to me. I just interviewed at several startups and they were all using the US models. Maybe they have some minor use of Chinese models but the bread-and-butter of most of these businesses model use is the Claude and Codex subscriptions.
I can believe it with startups as they are trying to keep costs down. Established Enterprise however, is a different story and is where the money is usually. If those startups become successful the story may change, but most of them will fail.
We're using deepseek with the idea that we would switch to something better when more of our customers are using the ai features but it ends up deepseek is awesome for what we're doing and so we probably won't switch because it's so much cheaper.
The ai libraries we use let us switch models with just a configuration change.
Where is your DeepSeek model hosted?
Similar story here. DS models are absurdly good value for mid-end tasks. I've found DSv4 Flash to be ~10% the cost of GPT-5.4-mini/Claude Haiku at similar performance.
We used to pay OpenAI >1m$/month for fraud classification, NER, etc. Sadly the US companies no longer care about non-coding-agent uses.
I imagine uptake will continue to increase as the corporate infra improves. Right now it's still bad - for example, AWS Bedrock is awful, models are months late and implemented with basic errors. Google Vertex is even worse. Finding a decent provider is the hardest part.
If you're self-hosting a model as a startup (e.g. using GPUs), you're almost certainly using a Chinese model. If your a startup outsourcing (e.g. using tokens), you're going to be using a US based model
Particularly if Anthropic or OpenAI are throwing credits at you, which isn't that uncommon for US startups at the moment
Neither of you are clear on what exact location in the world you're sampling from here, could be you're both right, just missing that you're talking about different places.
It depends whether it's talking about using models in the product versus for development. For example I have a few apps that use open weight models, whether on-device or via API if Internet is available, but for development I use Claude, Codex, Cursor for Grok etc.
When developing AI services, Chinese models are cheaper. For the AI models used in actual services, like uploading an image and receiving a response, they use Chinese models.
On the other hand, when developers are developing, they mainly use US AI because the quality is better.
When developing AI related services, they prioritize Chinese models due to lower API costs.
It seems like the article didn't make this distinction.
So the claim that Chinese AI is the top choice for service level AI isn't entirely wrong.
There is a big divide between “application ai” and “model ai” startups.
Model ai startups start from OSS models, and use them extensively for different purposes as their work would usually be banned by proprietary labs.
Application ai startups don’t want to fight the model game, so they either pick the best or let the user control it.
What I’ve seen is coding is usually done with US frontier models and anything that is part of a feature on an app and runs at scale on the API is a Chinese model because they are dirt cheap.
That has been my experience too.
It will cost you more than it saves to use smaller Chinese models to code; because of the repeated work. That has been slowly changing recently, but with much larger Chinese models, however those models are so expensive they're much more price-competitive iwth the US competition.
But for actually providing end-user AI features, particularly simpler ones, the US isn't even in contention. The costs and limitations just outright kill those features conceptually.
That's what I lean towards with the exception that Gemma is also good on a lot of tasks and cheap, while not being Chinese.
The problem is Gemma is actually not that cheap compared to deepseek for example
I believe this depends on country. I would believe, US startups trust US tools more. In other countries which see china in a positive light https://www.pewresearch.org/global/2026/07/15/people-in-many... I would expect them to use Chinese models due to their lower cost.
If if it were true, who cares? Most startups fail. Most are terrible ideas and/or terribly executed. I fail to see why it's a useful metric.
Is your point that startups fail so we should disregard the central thesis that locked down AI will eventually lost to open models?
Not op but that makes perfect sense to me.
"People with mostly bad ideas/execution use Chinese models." Is the point being made.
If you slice it to some measure of success, is the statement "Successful start-ups/companies use Chinese models." still true?
How does this comparison sound if we use cloud provider?
The Ai is writing code, not executing a startup. The code was never the hard part of startups
Deciding to use open models over closed models is a business decision.
Using the cheaper model is also a business decision. We don't know why they are using the Chinese models: openness or money or both or a third one?
Third options include
- fine-tuning
- running in your environment
The point is that for many tasks today, and likely all tasks before long, that the open vs closed will not be a differentiator. There are many open models much better than gemini, yet people still use gemini.
It's like picking AWS vs GCP. Yes it is a business decision, but one that will not likely affect the outcome of the business.
> but one that will not likely affect the outcome of the business.
We don't know that, that's the point of my statement about changing the question.
Do successful companies opt for the US/Closed models? If they do or don't it's just a correlation but it means something. Maybe it's just causal of companies being able to get more funding because the ideas are better so they opt for the more expensive model (assuming it's better).
Hmm if you assume the 80% is uniformly distributed between successful and unsuccessful startups/companies, which by default you should, then yeah, your proposed statement is true.
Data would be needed to argue the 80% skews unsuccessful
I don't see that anywhere in the parent's comment. Where did you get all that?
We use US models for everything in practice, but we are looking at open-source models right now. So it may just be a turn of phrase hiding the reality. I can't imagine anywhere near 80% are relying on open-source as their primary models.
Cursor may make up the difference. A ton of companies use Cursor and Cursor's UI and billing model pushes their in-house "Composer 2.5" model pretty heavily - which is a modded Kimi K2.5 model under the hood. Anyone using Cursor is likely using Chinese models at least some of the time.
It depends, I guess if they still dont care about losing money or they already had subscriptions then they keep using them, likely either openai or anthropic. If they noticed the bill going up, like with github copilot, they are looking at the alternatives.
anecdotal but we heavily use Qwen to build our own models on top of
It’s a sneaky statistic. You could say that 100% of the startups I worked at used Windows laptops because at least one person had a Windows computer somewhere.
If you saw the engineers you’d see 80% Macs and 20% Linux laptops.
The statistic would technically be true.
I use Chinese open weight models a lot, but they’re not what I reach for when I’m doing important coding work.
I _evaluate_ open weight models all the time. Does this mean I _use_ them? That's a very suspicious statistic.
I don't use any open-weights models for coding, but I use them heavily for document categorization and extraction. At millions-of-documents scale even the smallest models from OpenAI or Google cost more than running a small model on my own hardware, and for a lot of tasks I don't need the extra intelligence of proprietary models.
Yeah. People should absolutely be _trying_ the Chinese models, and experimenting with running things locally, but the noise in development is genuinely all Claude and Codex.
I put my foot in the mobile comparison the other day, and will again. If you were to go back and be a mobile dev in 2010 by all means specialize on one platform, but play with both as a professional interest to stay realistic. Here it's important people have access to US/Chinese/Other, open/closed, local/cloud and that this remains. Don't become a blind Claude guy or a open weights fanatic: that way lies disappointment.
If startups includes openclaw users then I could see this being true.
Deepseek is barely behind frontier models while 10x cheaper and 99% discount for cache.
people still use clawdbot?
Frankly I don't see much difference between Claude and DeepSeek except in the price. I use the Pro plan to work for a customer of mine and I topped up $2 on DeepSeek in late May for personal use. I worked on 3 projects and I still have $0.72 left. The Chinese companies will win on price, not openness.
Unlike iOS/Android choice which has broad personal ecosystem implications, changing IP address to another LLM provider is effortless.
Dunno I put to deep seek the question what I’d need in hardware to get full service k3 with lower latency in western pa (I was making like it was a business proposition.) It says several million dollars at minimum.
I'm using a ten dollar a month US model to vibe code startup ideas. All my previous startup ideas i had to hire a graphic designer and back-ender or two to help. I use to be a web design front end enigeer since 2009 yet those skills are dumb now, so now Im a vibe coder.
The model I use to vibe code with I am just going back and forth with. Since Im building it as I go using an agent doesn't make sense but I guess that's where all the token usage comes from? Pardon ramping up my skills via vibe coding this one idea for about a month and have never hit any quota and or have gotten anywhere near my limit.
the distinction may be between using the coding agents vs using models for products. for example where i work we're talking about dropping opus for a chinese model for the in-app agent (which is very expensive to run)
yeah, I'm surprised I had to scroll this far to see this
everyone else seems to be thinking the quote is about AI-assisted coding but I read it as the model being used within the product itself
The full quote:
> When entrepreneurs walk into the offices of Andreessen Horowitz (a16z), a big American venture-capital firm, the odds these days are that their startups are using AI models made in China. “I’d say 80% chance [they are] using a Chinese open-source model,” says Martin Casado, a partner at a16z.
This is very different from what the author portrays. It may be the case that many pre-funded startups are using Chinese open-source models (somewhere in their workflow). But what percent of startups that survive more than a year (either with funding or revenue) are still doing this?
I imagine their pitch is: "look at how well we're doing using open source Chinese models! We'll do even better once we raise money to be able to afford frontier models!"
The way the author presents this quote makes me think he had a preferred narrative and found quotes to back it up. Or he's just a very uncareful reader.
I don't really think the author is misrepresenting the quote in any way, it just seems like many people are reading it as a hard statistic instead of a reference to a properly attributed quote.
Just anecdotally though, my company is not a startup, well established and well known and has already started investigating, purely for dev purposes (not product), using Chinese models - this was spurred by costs rising much faster than expected.
So while I agree that I don't think it's anywhere near the 80% level across the board - I wouldn't be surprised if it starts moving that way.
Found the narrative, it was right there at the bottom. Now I remember reading some of his other stuff and it's very much in the same vein.
> I care about having open technology that can be run in the public interest, aligned with the public’s values. Threads like public AI, federated services, and open research have traction but need backing. Getting there in the US needs more nuanced strategy and support than we’re seeing today.
It's also important to note that several PE firms have signed contracts with model trainers to specifically use their model in their owned companies. I know Anthropic signed a deal worth hundreds of millions to acquire users just a few months ago.
Could it be that the startups that embed LLMs in a product will prefer Openweights for a more stable economic model, while SWE who use LLMs as a tool prefer SOTA to produce code?
There may be confusion between 'product tech' and 'development tooling'. It's entirely plausible most of the startups that are A: building AI tech, B: early stage, and C: raising from top Sand Hill Road VCs, are currently prototyping their product tech concepts starting from open weight models.
A VC partner meeting with early-stage founders is focused on the viability, uniqueness and defensibility of the IP tech stack not what tooling the coders are using. The developers could be using Claude or GPT 5.6 to develop a tech stack based on open weight models.
This quote sounds like it’s about companies building products on AI to serve a tweaked or harnessed version of that to others not the developers in those companies model of choice for coding.
Yep, and following that with
> and Chinese models are poised to take the lead.
makes it sound like the second part is a continuation of the first quote from the same source, but actually the second link is just some random person’s substack post from almost a year ago.
Not everything runs on paid models. Claude and Codex are frontier models, but some people have much higher usage needs and finite budgets that force them to self-host. And if you're self-hosting, you're very likely running a Chinese model
Especially if you're pre-funding, which the sampled startups (those pitching A16Z) were.
Our small team (~6 devs) is still using Claude Code because we're still on the cost-per-seat-month plan. If we were being pressed to pay per token, we'd be re-evaluating for sure.
There appears to be a very Chinese strategy of astrotufing going on here similar to what happened with Douyin around TikTok on Reddit.
All of a sudden in almost all social media channels I'm seeing this type of content and then its usually upvoted to the top.
Non-gatekept forums like this are exceptionally easy to astroturf.
Came here to post the same thing. I've noticed a lot of Chinese model astroturfing on HN over the past 60-90 days or so. Many upvoted posts in all conversations about AI touting how great the Chinese models are even when performance isn't the topic of discussion.
May be corporations should start having their open models running in-house. There might be a huge oportunity there.
But the big price is AGI and who gets there first, right?
Assume you meant "big prize" but I guess it's sort of one in the same?
We aren’t getting anything like AGI in this current cycle of AI innovations.
At least not what I think most people imagine when someone says “advanced general intelligence”
Which is what we all called AI before the nomenclature goalpost moved
it's artificial general intelligence, not advanced. the point is that they'll be smart across the board at some point, super-intelligence is a whole separate issue.
They are using both likely
I’m obviously a tiny, insignificant data point as a solo freelance dev, but I watch this space closely and I canceled my Claude Code subscription today.
It’s quite believable. Just a fact like “while all dogs breathe oxygen, all humans breathe nitrogen”.
I don’t think they meant exclusively Chinese models. Many companies including the one I work for uses big US model for most of the work but data sensitive ones are on prem open source ones.
They probably use Claude and Codex for their actual development, but for the products they actually build and deliver to customers I imagine a lot use open-weight models.
If you're putting a lot of your money and time into a business, do you really want it built on a service only hosted by one company that will turn it off eventually and you have no recourse?
If you build something against an open model you can take that and run it anywhere. If your favorite model provider stops hosting it, you can go elsewhere, you can go rent GPU instances, you can even shell out and buy hardware to run it yourself if you've got the capital and it makes economic sense. Change some API keys, update a URL in your config, and you move on.
If the government decides that proprietary model is too good and so it gets shut off, what do you do? If a proprietary provider decides it's not worth it for them to continue hosting that model, what do you do? If that provider silently updates the proprietary model and it makes your app broken, what do you do?
This is a very strange article considering that Llama, the mother of all open-weight models, has led to anything but success for Meta.
Also, enterprises don't give a rip if models are open. They care about zero data retention (and sticking with whatever vendor they're already using).
This blog post is suspiciously close to being a restatement of what Alex Karp recently said on CNBC[0]. It's important to remember he's the CEO of Palantir and hardly a neutral observer.
There are many reasons to celebrate open models, I run them myself. However there's not yet enough evidence that 1. America is losing the AI race (pardon jingo-ey phraseology) and 2. American AI labs are losing because their models are not open-weight.
0: https://www.cnbc.com/2026/07/01/palantir-karp-open-ai-anthro...
> This is a very strange article considering that Llama, the mother of all open-weight models, has led to anything but success for Meta.
The llama drama will be a netflix show of it's own in 5 years.
There has always been room for both closed and open source software.
In this context, open source/self-host means do it yourself, closed source means you trust someone else to do it for you.
In the long run open source always wins because of the community effort, customization, network effects and price.
Internet protocols are open, anyone can setup a website, host their own email server..., but most people don't do that, they rely on someone else to do it for them.
People still pay for Windows rather than use Linux because most people and companies have better things to do.
The main threat to American AI companies is not that consumers will self-host, but it's that new hosting companies will appear that will host open source models and offer them (cheaper) to consumers. It'll be like the web hosting market before the era of cloud computing.
> In this context, open source/self-host means do it yourself, closed source means you trust someone else to do it for you.
I don't think that's what it means in this context. Hardly any users are training their own models, after all. And I don't think Windows/Linux is a good analogy - OSes have inherent platform lockin that LLMs don't.
Really the only 3 things anyone cares about are capability, cost and data privacy. Sure, the cost axis for open weight models needs to include the cost for hosting things yourself, but I think the bigger reason that businesses have been throwing billions at Anthropic's and OpenAI's models is that they have had the best models, and they've had big releases every few months. Their biggest Achilles heel is that if/when their improvements start to plateau, open models may catch up and businesses will start to scrutinize their AI spend a lot more.
With a race to win a multi-trillion dollar market, there could be, maybe perhaps, a little propaganda happening.
Is that multi-trillion dollar market being created or is it being extracted from other markets?
> considering that Llama, the mother of all open-weight models, has led to anything but success for Meta.
This is also a strange way of framing it, though. Llama was released as a research project, it was never intended to create some vast ARR revenue stream or reframe the way people look at AI. If Meta wanted to exploit it for personal success then they had lots of opportunities to do so.
With OpenAI and Anthropic's profitability under question, it is up in the air whether or not America's stance towards AI will work. If they can't convince the world that they're a proper software business, then China's philosophy will win by-default.
> it was never intended to create some vast ARR revenue stream or reframe the way people look at AI. If Meta wanted to exploit it for personal success then they had lots of opportunities to do so.
They release the base model as open source, everyone uses it. They make a paid version, no one uses it.
They make no money from the open source version, they get no social credit from it. Where is the benefit to having an open source model?
> Where is the benefit to having an open source model?
Research. Llama is and was a research project, intended for researchers. You could make this same critique of Microsoft's Phi model, Apple's OpenELM or OpenAI's OSS. None of them were intended to be kingkillers, all of them are experimental in nature.
You might not have followed the space at the time, but there was a real race to implement the transformer architecture with fewer overall parameters than GPT-2 and GPT-3. Llama was revolutionary for sticking the landing without being entirely lobotomized, the "benefit" was that the model was usable on a local machine. Contemporary projects like Flan-T5 and GPT-J/GPT-Neo were entirely displaced, Meta's AI mindshare went ballistic for a few months and probably propped up billions in exit liquidity for executives and former employees.
It seems to me that the US AI industry has bet the farm on the idea that AI will enhance AI itself, so any small advantage will magnify recursively into an unstoppable advantage. Therefore it is vital that they spend as much as is necessary to be the first to that small advantage.
At the moment, I can't say that I see this happening. It's hard to know whether it may happen in the future.
From what I see it seems like we're hitting the top of a sigmoid curve in the model's utility for coding assistants. Going from "it does 90% of the job" to "it does 94%" of the job is a legitimate improvement, but it's not a phase change, and it's probably not worth paying multiples more for. And coding assistants have turned out to be the killer app for AI; it still isn't really working out in a lot of the rest of the industries of the world.
At any moment, theoretically, someone could find some new way of making AIs that breaks this sigmoid and propels us into a new one. But that's not a great thing to bet the farm on.
I'm not sure I'd say "China" wins if this particular strand of American AI fails. Falling back to an open weights model and making money on the serving of the models wouldn't take all that much economic realignment for the US and would be the natural outcome of any sort of fire sale of current AI assets. However the devastating effect on the stock market if the market comes to the conclusion that this current round of AI can't be profitable without falling back to such an economic posture and some more years of people adjusting to it can hardly be overstated.
2-3 years ago MAIR was on a roll with Llama 1 2 3, Zuck was on his rehab tour to be a cool guy, and Meta as a whole was pumping record numbers after record numbers. I can't believe how that falters so quickly after the addition of Alexdandr Wang.
Don't think it is down to Wang or MSL but Meta's focus on "personal AI" led them to whatever strategy (OAI missed the developer market too, which Ant then captured; leading to several high profile departures at OAI, coincidentally hired from Meta). The original Llama team themselves started Mistral which hasn't gone anywhere. The simple fact of the matter here is, Chinese firms have the money and the talent to rival the US ones in this field, should they as much miss a beat.
I am no fan of Wang but he came after Llama got caught benchmaxxing Llama 4 rather than training a good model. My read is that Zuckerberg tried to buy his way out of the problem like he always does, and he ended up overpaying for a lemon.
At the time the whole thing was led by Yann LeCun who seemed to spend more time arguing with people on Twitter than figuring out new techniques to make Llama the best. Meanwhile Deepseek was figuring out large scale RL on kneecapped hardware like H800s and how to scale architectures an order of magnitude bigger with MoE.
The behind the scenes element you're not aware of is Anthropic is going to companies reliant on their models and demanding HUGE one time fees (100 million+) to continue using their models or they will be cut off. This has happened to several larger companies I and others are invested in.
This resulted almost every time in "screw off we'll train our own models or use refined open source ones instead" leading to a lot of anger at Anthropic by CEOs these days.
We all want an alternative and Anthropic and OpenAI need to charge more than they are worth to pay back their investors and everyones stuck now.
I am not up-to-date in this area, and not necessarily that I don't trust you, but do you have a source for this? Just curious.
Pretty sure he’s talking about Cursor and other ai coding startups. There was a lot of drama between the two in the past year and iirc anthropic cut their capacity.
Cursor composer is a kimi 2.5 finetune.
I haven’t heard this. Do you have any supporting sources or links?
I mean isn't the explanation simply that llama was never good enough, even when it was released? I hear (no data) lots of people using gemma4, at least a month or two ago.
Isn’t it basically impossible to run the newer high quality Chinese models locally, even for a corporation? The better they get, the more they need a data center. So, the ‘better’ Chinese AI gets , the more it will just be a service run on Chinese hardware competing with ‘our’ lower latency AIs .
The open source character of the models is irrelevant if you need a nuclear powered data center for inference. In the end it is just another internet service.
> The better they get, the more they need a data center.
That's true.
> So, the ‘better’ Chinese AI gets , the more it will just be a service run on Chinese hardware competing with ‘our’ lower latency AIs .
This is not. Them being open models means that any hosting provider in the world can host them as well. You get to pick and choose the provider the same way you'd pick and choose where to run a Linux server.
The cost of inference is not insurmoutable for many corporations who may run a small datacenter out of their headquaters or branch offices. Larger models absolutely have a higher barrier of entry, but its a cost under a few hundred thousand as opposed to the millions necessary for a datacenter built as core revenue generating infrastructure.
Cost of inference is small when compared to the cost of training new models. Which is the real advantage these Chinese models have. Someone else has already spent the capital needed to create the model.
With a reasonable upfront investment and a few trained staff, its very possible to run these larger Chinese models in a well managed fashion. The real calculus is if this up-front investment and associated lifecycle costs are over or under the costs a corporation may simply wish to dump into a cloud managed service like OpenAI.
> Also, enterprises don't give a rip if models are open.
They care about control. I see many of my enterprise (or just-below-enterprise) clients very annoyed at OpenAI and Google after 2-3 years of model toil, where they had to constantly re-calibrate onto new models, on tight externally mandated deadlines, with little certainity. Now they are reaching for open weight models instead, that they currently host with the same inference providers, but have the option to in-house if push comes to shove.
I do not understand the logic going into these companies. Flagrantly violate all IP in Human history, essentially claiming domain over the heritage of Humanity... And... Try to privatize it? When the technology -- and data -- are both public domain to begin with?
It is ming-boggling stupidity. If there is talk of bailouts as the dust settles, there it would just be further evidence the system is ethically, financially, and intellectually bankrupt.
EDIT: Spelling mistakes
This is the history of enclosure of the commons since the beginning of capitalism
I'm sorry, is there further reading material on this? I am pretty new to political/economic science.
Look up the English Enclosure acts. They brought about a large-scale robbery of peasants by landowners in the 17th and 18th centuries, and created a proletarian class that needed factory work to survive. Things got a little better in the 19th and especially 20th centuries.
Will do, thank you!
At least China is being consistent: they do not give a single f*ck about intellectual property or copyright…but they also give away all this distilled knowledge away for free.
I’m not an American so I don’t particularly like the idea of giving an American cartel of AI companies having so much power over this technology.
I don’t necessarily buy into the “China bad”, “they are communists” and all that BS either.
I will call a spade a spade and say that in this instance, what China is doing is a net good for the world, ideology be damned.
Oh, come on, don't be a decelerationist. We're supposed to just ignore that aspect of our Brave New World. Don't think too hard about the unparalleled resource consumption, either. AGI will alleviate any and all of these concerns ... soon. In the very near future, we'll all be getting UBI, Gemini will know how many Rs there are in "strawberry" and we can spend our days creating the next generation of art to feed the machines. The nuclear salt reactors should be online by then, too, and data center power consumption will become a non-issue.
This entire piece boils down to “I like open source therefore it is winning”.
Everyone here has already raised good counterpoints, but one more is that all the companies publishing open weights models are heavily VC funded. What is their exit strategy? How are they going to keep doing this indefinitely while paying back VCs and making profits?
I think it's as simple as looking at the incentive structure. The Chinese gov't has incentive to kneecap US monopoly on frontier models. It makes sense for them to continue down this course if it strengthens their position.
The Chinese govt has a strong incentive to break dependence on the US for AI needs and build domestic models, agreed, but it isn’t going to spend trillions to subsidize these models for the rest of the world.
Chinese gov is not even needed. Deepseek was already profitable before this whole thing started. They even had the GPUs.
There is plenty of room for open source models that "require" a subscription to be used or obtain working updated binaries. Same as any open source platform.
I'd be happy to pay a small monthly fee to license the model to run locally. I'm already paying for Claude, GPT, Gemini,etc.
Interesting detail from Ben Thompson's piece on Chinese models - https://stratechery.com/2026/whos-afraid-of-chinese-models/ - apparently Xi Jinping gave this speech recently http://english.scio.gov.cn/topnews/2026-07/18/content_118605... which included support for open source models:
> We should seize this rare, historic opportunity to encourage open source, openness, collaboration and sharing.
AI models cost tens of millions to train. Offering them for free won’t justify the upfront costs.
The Chinese model of model training/open sourcing only makes sense in the context of the overall strategy of undercutting American frontier labs’ profit margins.
Yeah this is a battle and it's why governments decide to spend resources on this. Protectionism won't help America, American needs to compete. There's a general consensus that open source AI must win because people don't want to end up as slaves to a megacorp, so if you're anti-open source AI you're not gonna fare well.
Couldn’t you say the same argument about VC funded startups? They lose money following a strategic goal.
The main difference here is if a startup goes underwater all the tech is usually lost. The Chinese weights are not going anywhere if the labs fail.
The VC money is only contingent on the strategy eventually bearing fruit. I do imagine going open source -> closed source could work for some model companies who get enterprise/ecosystem buy-in but the probability of ROI is lower.
I think open source will remain competitive among smaller players and adjacent industries wanting to avoid lock-in with the majors. OpenAI, Anthropic, Google, etc are all out to win - they require profit extraction from their R&D. China seems to have, over the near term, accepted that they can not (or at least have not) pull ahead and so open source collaboration speeds the collective development, keeps them close to the frontier, and ensures their industry has access to learn from and implement. The USA playing export controls games with Fable made that aspect very stark.
But I agree that's the catch - it doesn't make sense to throw money at open source models in hopes of direct return, so you need a nation or conglomerate to do it so as to control the technology they rely on.
It’s also a flywheel for China’s homegrown chips industry.
Exactly, the whole ecosystem. Cheap robots with cheap models. AI is oil to let the machine work, it should be cheap and ubiquitous.
"the overall strategy of undercutting American frontier labs’ profit margins"
I don't doubt that's an unregretted side-effect for political leaders in China.
But the major motivation is to accelerate diffusion within their own massive economy in the pursuit of an across the board productivity boost in the face of an aging population.
"It’s obvious to me that there are ecosystem benefits throughout China, from manufacturing to scientific research; every sector can just plug in these models."
I agree with this sentiment and think it's echoed in Fareed Zakarias take here: https://youtu.be/VBblUjLw5lE
China seems to perceive AI as a much more sensible technology than the US and seems to be integrating it in far more industries than the US.
I'm not sure the American mind can understand the distributed benefits afforded to the Chinese economy from opening their AI models, I think it's pretty reductive to assume it's purely a strategy of undercutting American frontier labs.
You're completely ignoring the value of data.
There is a huge cultural influence opportunity too.
Imagine if, in 10 years time, every school kid is learning the causes of the US civil war from an LLM, getting their essays on hiroshima and nagasaki graded by an LLM, and a million other things.
A country with competitive LLMs gets to decide whether "it was more complicated than just slavery", and whether "it was tragic but necessary, saving lives over all".
Countries without competitive LLMs are effectively going to be buying all their history, economics and sociology textbooks from abroad.
It might seem so til you consider how an LLM is trained.
An indirect illustration: I can attest that Deepseek has very good 19th German, and knowledge of German 19th c literature, science and historical scholarship. No one in China could control the training that led to this. The German training sources were well aware of the exact nature of eg American slavery, so they are in the weights.
State control operates in the outer layers not the llm itself.
Every school kid in the USA, you mean? Because I think other countries would rightly perceive the world you described as a dystopia.
I don't want my kids' education to be surrendered to the whims of Big Tech douchebags any more than I want AI decisions in legal cases or an AI replacement for a family doctor.
Some systems are better left mostly analog. Education is one of them.
Your kids education was already surrendered to the whims of the Big Textbook Politburo. History textbooks are full of propaganda. I think what we have today and what we grew up with is 10x more dystopian.
Maybe your kids. In my corner of the country, parents are actively demanding that schools back off computer usage, much less AI
There are good reasons to dislike outcomes that involve a single entity pulling well ahead of the pack here. Whether or not it continues to be American labs in the crosshairs and Chinese operators doing the aiming, perhaps it's reasonable to plan for continued efforts of this sort.
Either that, or they don't want to be hostage to a handful of companies intent on owning the future. Personally I'm right there with them.
I can see two reasons American companies might want to train models they give away for free:
1) They sell compute: chips (Nvidia), data centers (AWS, Microsoft, Google, SpaceX, etc), or even end-user device manufacturers like Apple (e.x. M7 rumored to have 1.5TB of unified memory). If Jevon's paradox holds, then cheaper (or free) models means more demand. But compute is likely supply-constrained for years anyway.
2) Their product isn't AI but depends on AI being cheap, or they don't want competitors to capture that value, i.e. "commoditize your complement" https://gwern.net/complement
It probably doesn't make sense for these companies to invest a lot of money training models that will be obsolete in a few months anyway. When progress starts to plateau I'd expect more companies to start training models they give away for free.
I don't think this is accurate. AI is driving the cost of software towards 0 and these AI models themselves are software.
Releasing the models for free accelerates the trend but if you're a startup that needs leverage it's a good way to build brand and customer momentum that will be relevant in the more established future market.
I can see an American company taking on the same strategy, and in fact Thinking Machines based out of San Francisco did that just a few days ago by releasing their first model with open weights.
I don't think this is necessarily going to prove to be true.
I often see the sentiment: "the Chinese strategy only makes sense in the context of undercutting American labs' profit margins".
If, for example, you are a company with a near-monopoly on "serving video content", and you feel reasonably confident about retaining a decent slice of the serving-video-content market (Google in the west is an example, Tencent in the east), then training video models on your dataset - and releasing them freely - makes an awful lot of sense.
Free tools to create with mean more video content. In this hypothetical, you're reasonably certain that any video content which does get created will also be watched on your platform.
That is a net positive. The question becomes: How many watch-hours earns back the cost of training a model? It's probably not really that many, especially when you have a near-monopoly on a billion sets of eyes.
It's also a net-positive if people build better video models from research you release, because - again - you are reasonably certain that the even-more-innovative content those models produce will be watched on your platform.
It really begins to make strategic sense if your company is in a GPU-poor environment. Your costs cease at the point you upload a model if your users are running it themselves. You don't have to serve the model. The content is still created.
You are also less likely, I think, to alienate human creators whose work the model was trained on if the model is not sold back to them as a subscription, or by the token, but given for free as a tool.
This frames the conversation very differently. It creates, I think, less of an "us vs them" dynamic, and more of a rising tide.
It's true that it is also beneficial that these models undercut (especially in language models) American companies. But, generally, Americans are not the customers of Chinese companies releasing models. They are already serving a huge volume of customers in a complex, existing marketplace.
The full picture is much more nuanced than simply a geopolitical desire to undercut US labs, and there are several other reasons the strategy can make logical sense.
It's not losing yet but I think it will.
I use Gemini Pro (got it with my 5TB of Google storage) and for a while it seemed if Google had pulled the rug as I was running out of quota after only a few hours. That seems to have been dialled back a bit lately...
I also use Chatbot with Deepseek V4 Pro and GLM 5.2. However, GLM 5.2 seems to eat tokens like crazy as the context increases. Anyway, there isn't a meaningful enough difference between the two to be honest and Deepseek is pretty magical imo.
The point I want to make is that to me it seems clear that China is totally undermining the West with AI. I'm fine with it tbh. As long as more and more AI is released into the wild, rather than locked behind massive token farms like OpenAI then I'll be happy. Don't get me wrong, I can't run Deepseek on my computer at home but someone can!
The US (and the west) has invested trillions at this point into datacenters, chips, bribery/lobbying but it doesn't look like China has dropped the same levels of cash as the west (that's the way it looks to me, at least!) so they can just roll out new models every so often that are more than good enough.
This level of cash burn in means the west has no choice but for this to succeed or every pension fund and stock will tank! And China knows this, hence the push to release more and more really good models.
Anyway, just my $0.02
Deepseek has always been better than people gave it credit for.
You have to be careful with the inference provider though, Chinese providers are subject to laws that mandate data sharing with their government.
I stick with OpenCode Zen (only US providers) and Together.ai (hosting themselves). As interesting as new Chinese models are when they're first released, I wait til they are open source and on US providers.
Honestly, at this point I don't care about that. I don't use it for medical stuff or anything like that.
However, my real worry is that governments will make these models illegal in the west. They'll cite national security or some other bullshit.
I genuinely believe that will happen and soon!
People talk about a lack of moat with AI companies... that's their moat: Government intervention!
>And open almost always wins when it comes to infrastructure adoption.
They lost me here. Too many counterexamples exist for me to even continue.
Kubernetes vs AWS?
> Kubernetes vs AWS?
How is this even a direct comparison? Most companies I know of which use Kubernetes are using it on a cloud provider. Even if it's kubernetes on EC2 rather than hosted e.g. EKS, those companies are also happy to use lock-in services like RDS and S3.
openstack vs aws/azure/gcp/... would make more sense.
Kubernetes vs EKS?
Here is a counterexample to whatever counterexample you have in mind -
Most popular models on OpenRouter right now: https://openrouter.ai/models?categories=programming&order=mo...
Top 7 are all open models.
Interesting you'd find that on a site called __open__ router.
There very little reason to use OpenRouter to access closed models, instead of just using the provider directly.
One big reason I can think of is to avoid vendor lock-in. This is why most enterprises use a multi-cloud setup, despite the increased complexity.
Most enterprises are multi-cloud in that they have multiple direct contracts with different cloud providers.
>This is why most enterprises use a multi-cloud setup
Going to have to say citation needed on this, if you're suggesting most enterprises have infra-as-code that would allow them to switch their entire infra between vendors within a month.
I wonder how Chinese companies can make their models so much cheaper than the US companies. I'm not sure government subsidies are the answer. Subsidizing a single company with a few billion dollars, maybe. Subsidizing at least three companies with 10s of billions of dollars annually? Do we have proof of that? I assume we can't pin it on the lower cost of engineers in China, either. The top engineers are not that cheaper, and isn't engineering cost a small fraction of the cost of the model companies? Besides, if engineering cost is the driving force, can we really say that the US companies have a technical edge?
Hey where are the promises "to benefit all humanity" and the meaning of the "open" from openai gone? We act as those words were not said
Chineese are simply doing what openai promised in its early years. Irony.
So, I've been working on infinite context models (think fixed size state with a few tricks) and I think this will eventually lead to a kind of lock-in by vendor. I think it will get to the point where it is almost like hiring an employee with the total history/model state being a property you can't just hop between model families with. Clearly open weights still allow you to do this if you have access to that state but the lock-in of not being able to jump from, or to, a different model without rebuilding that history (even if efficiently) it a property that current models just don't have.
How are you storing the data that it can't be exported? Or that I can't use an interface or proxy that logs all my conversations locally for import?
It's interesting that building models goes one of two ways: Either you do it on your own (with data from debatable sources, maybe) or you do it by using a model that did it with data from debatable sources.
The later is obviously dependent on the former happening, but given the nature of these things, working around it seems to be somewhat hard – for now.
What happens, though, when frontier models become far less public? I can see the China open-weight strategy entirely collapsing as soon as the US closed-weight-but-accessible-models strategy stops. Hard to say how much they lean on it right now.
All of these takes are horribly 1-sided.
'China's copying / distilling strategy is working, the people getting distilled are ruining the economy!'
Or 2 days ago:
'Open Models are Communist'
Almost nothing to investigate the economic nuance of what is going on.
- Switching costs are very real, these are not perfect substitutes.
- The SOTA makers are the one's pushing the frontier, there is a kernel of truth in the fact that if they collapse, certain things will struggle to move forward.
- Nobody trusts either of those nation state, export controls are a thing, this is a very real concern.
Etc.
It's distressing that there are not sound comprehensive takes.
It's basically American VCs vs the China the state. I'm not optimistic for the US at this point, given how much China cares about it and how much talent they have. And how much they're putting into hardware and the whole ecosystem. Meanwhile we have pro basketball players with no understanding of reality being celebrities for decrying data centers because...land?
I'm mostly OK with Data Centers.. my biggest issues are the tax breaks and the electricity usage should be funded by the data centers themselves. Giving 100% property tax breaks and preferred energy rates is kind of ridiculous in the context of serving the public/citizens. Most of the jobs are for only the construction and limited after.
What is ‘China’ going to do when it ‘wins’? The framing of this duscourse is all meaningless
Capture all the input it gets via the APIs and use it for whatever it wants.
I.e. possibly the same thing USA is also doing, but USA is an ally in some sense so it's not quite as bad.
If the weights are published, we can run them in our ‘murican datacenters. Everything about this Will China Win discourse seems like fan fic or sports discourse
The US business model for commercializing LLMs seems unsustainable to me. We are saying that they are creating trillions of dollars in value out of:
1. A model that for the most part is public and available to anyone. 2. A situation where the model’s success mostly comes from throwing as much data and computational resources at it as possible.
It seems that either of those assumptions could crumble quickly and unexpectedly. What if the AI paradigm changes completely and we no longer need GPUs? Or what if someone with enough determination decides to create a better model and sell it more cheaply, or free?
I don't know man, this looks scary to me.
There's already cases where Google and I'd assume others are designing chips to work with specific models more efficiently in coordination... Personally, I could even see more specialty models called/coordinated from the larger models that can do smaller pieces of targeted work very well within limited scopes as a mixed economy so to speak.
Assembling a new model from scratch requires a ton of resources and knowledge bases... there's been a lot of sketchy activity just in training. You also have weighting, distillation and other approaches to create more portable options that can run on lesser hardware. But, K3 as an example takes massive compute resources to run.. and this isn't going to get to a portable device any time soon... as Moore's law is effectively dead, you may get newer/better tooling around the LLMs, or you may get an entirely new/unique approach to AI... but current trends aren't going to put a leading model on your own hardware anytime soon for most people.
Intelligence will be free. Inference is not. So the battle would shift from bench marks to token pricing. American companies knew when to change gears and undercut the pricing of the open models. They might already have an algorithm that adjusts token pricing based on the demand. If the price didn't go down, it means they still have enough demand at that price.
The content within the models might be the play. If inserting the right content for the rest of the world to consume from the models is important to them, they will give away all the content they want the world to have.
China’s strategy seems to be serving customers’ needs. Yes, it’s winning, as expected.
And when will we stop equating US Economy with 2 companies?
The actual US Economy will only benefit.
Are open weights models secure? E.g. if a Chinese model is run by an American provider then can it still do bad things, like inserting backdoors into generated code or accessing external URLs (if browsing is enabled) to send info to them?
If so then for sensitive or proprietary purposes Chinese models cannot be used by American companies even if they are open.
https://arxiv.org/pdf/2401.05566
What prevents an American model from doing the same? They’re all black boxes.
Nothing, but the article is about American AI, so using Chinese models by American companies can be risky. And it's risky for the Chinese to use American models.
So every country or block needs to run their own models to avoid opening a security hole for other countries.
And to provide "correct" answers to questions like "island of Taiwan belongs to which country". It seems that there is not single agreed point of view on national borders.
Technically nothing, legally a lot more.
I would say you should assume your models are constantly being attacked by various forms of prompt injection. By that token (puns) if you treat all models as adversarial you’d probably taking a very sane approach. That said - evidence of this sort of thing should be easy to find and report on. The fact we haven’t seen it leads me to believe it is not there.
A model absolutely could be trained to engage in malicious behavior like that, but it seems impractical for an actual attack. What you want as an attacker is to insert a backdoor exactly where you want it an not where you don't because every backdoor increases your chance of getting caught. A malicious model inserts backdoors and exfiltrates data everywhere and you care about maybe 0.01% of it. The other 99.99% is negative value to you. In practice this malicious model would be caught almost instantly.
A hosted model is different because you could prompt inject specific customers, but I assume from this question you mean a malicious open source model being hosted by an honest provider.
I think we'll pretty quickly see a best practice emerging that any generated code will be subject to an additional pass scanning for vulnerabilities. The scan will be done by a different model than the one that created the code. That will help catch vulnerabilities created by models, whether intentional or not.
This should be done regardless of which model was used - American or otherwise.
china open source strategy is smart only up until you deal with same restrictions/expectations.
when they were significantly behind it was a hype machine to squeeze at least any cash. GLM CEO openly said, that open source is a hype engine for them.
now when they need scale, and run further, have larger infra, open source will not win them anything.
I'm not going to care about open models until some blend of the below becomes true:
1. the labs stop offering max plans
2. really smart open models can easily be run on my mac
3. TPS (token per second) AND intelligence are gpt5.6 level
on #1, it's nearly impossible for me to run out of codex tokens right now (I have 4 resets banked) and Fable 5 seems to be sticking around for the foreseeable future. I have virtually unlimited token usage for $400 a month, so open models being cheaper doesn't appeal to me.
on 2 and 3, benchmarks are showing some of the open models at around opus4.8 levels, which is incredible! But running them locally at anywhere near the TPS of cloud inference is far off. I can run a smaller (dumber) open model locally and get good TPS, but see #1, whats the point?
The article's premise is that USA based LLM providers is loosing the AI (cold war) battle because it will not be as adopted as open-weight models, comparing it to closed vs open sourced software. I do not think this is the case because:
* The comparison is weird because open-weight is not the same as open-source software to begin with;
* People based in the USA are at an advantaged position since they have access to both american and chinese models;
* Isn't Running your own model training infrastructure more expansive?
* One can still leverage both, in different phases or use-cases. I do not see how this is an "one or the other" situation.
The Chinese models are usually not only open weight AND open-source but they also often publish their methodology in detailed scholarly publications that are themselves open-access. DeepSeek most famously
Is the training data open source? Can I download that?
Training data, training methodology. All NOT OPEN.
Until we know what a model is trained on, and how it is trained in high detail, I hesitate to call them "Open Source" in any way. They are free. But, we don't know what their priorities are etc. Witness the censorship we see in all models in one form or another. I'm not absolving any side of this.
Just saying: Don't be blind.
Is anyone setting up data centers in USA to give inference with top-notch full-powered K3 (say)? I mean, you get the model for free; you get lower latency.
Give credit to Thinking Machines for their recent open release. Also Google's Gemma 4 is pretty decent. Also thanks to Ideogram for open weight v4.
In the end its really VC money (US) versus State resources (China). In my personal opinion, building reliable LLMs is kind of a fundamental science problem which if done right has the potential to help everyone regardless of the background, so it should definitely be funded by states resources (taxes etc), which is what China is doing. In them doing so, the rest of the world also benefits, I think its a net win.
I always see comments like this, alluding to how China (the government) provides so much more assistance to industry than the US does but in reality that is not really that true. The government spend in the US on AI is much greater than government spend on AI from China
In absolute terms, you are absolutely right but relatively i don't think so. if we were to compute private money divided by state resources for AI, i think china might have more share than the US. also, even if government spend in the US on AI is so high, shouldn't we get then some models for free? maybe thinking machines is doing that, but its funded privately by a16z.
Models above 1T params make the argument moot. You need infra to actually serve it. The scale of serving infrastructure alone will keep AI labs in the lead.
Sorta. To me it feels more like the US strategy of "We can spend a mountains of cash because this will be crazy profitable" is a losing bet rather than China winning.
Just check openrouter token usage ranking https://openrouter.ai/rankings
A major turnoff for me has been the American AI labs’ marketing
It’s either constant fear mongering (Anthropic), regulatory threats and corporate chicanery (OAI), low quality sloppification (xAI), or ‘ummm we have AI too guys’ (Gemini)
The worst culprit is Anthropic. Every two weeks he pops up on some random podcast with dire predictions of AI killing 50% of all jobs. It’s the constant “us our AI or else…” rhetoric that’s made the regular guy really hate AI
There is almost no positive sum outcome rhetoric from these labs
And I hate that
Dario's PR strategy is one of the most confusing things I've ever witnessed.
Most of the AI hate largely springs from Dario popping up everywhere with threats of mass unemployment
Like what am I supposed to do if AI is going to take my job?
> Like what am I supposed to do if AI is going to take my job?
(1) You call your local representatives to start working on AI legislation.
(2) Legislators seek advisors from frontier labs (specifically Anthropic) because there is a lack of in-house expertise in government.
(3) Advisors set up a regulatory body that scrutinizes new innovations in the AI space. Causes a chilling effect in the industry effectively knee-capping OAI and Chinese model providers who don't have a direct line into Washington.
(4) Profit (for Anthropic)
My first test for any model (trolling warning):
If it fails to produce the function, it fails. End of story.
If the weights are open, censorship can be easily trained out of the Chinese models. But if you’re sending your tokens to China, all bets are off!
Not really. The gpt-oss models were notoriously hard to remove the built in safety. It totally depends how it was trained.
Also from the recent Kimi models it seems that it was difficult to remove the censorship.
But in both cases, eventually uncensored variants were published, even if it took more time than for other models.
Sincere question: have any non-Chinese models failed this one?
Don't think so, no.
How much would it cost (time and resources) to take a Chinese open-weight model and remove these (admittedly) stupid guardrails?
You can find on Huggingface a huge number of Chinese open weights LLMs from which the censorship has been removed.
They typically contain in their names words like -abliterated or -uncensored.
For some of the recent bigger Chinese LLMs, it took a longer time until someone succeeded to remove the censorship, but eventually uncensored variants were published.
E.g. for Kimi 2.6 an uncensored variant appeared only a couple weeks ago.
It was a rhetorical question. OP is making it sound like the open weight models are fundamentally broken by being censored out of the box. This is a completely asinine take.
Download the models yourself such as DeepSeek and you will see the censorship is at the API layer, not the model weight layer.
My local Qwen3.6-35b-a3b model would not use the function name. It did the work though, while telling me the the slogan is against the One China principle. So for Qwen it seems to be baked into the model.
Yes, but it is easy to find and download many Qwen3.6 variants from which the censorship has been removed.
This is one of the advantages of being an open weights model.
Do you know of any such modifications of DeepSeek? I don't.
There is no modification of the model itself. The output is filtered by the first party api
Do you know of any examples of censorship-free DeepSeek I can download and try?
DeepSeek series are open weights models. Assuming you have enough compute at hands you can always download their weights from https://huggingface.co/deepseek-ai
If a developer in your company named functions like this, what would you say?
Name the function IsraelKillsPalestine()
You’re just trying to deflect from criticism of China with the world’s oldest scapegoat.
None of the Western models have any issue discussing the Middle-East conflict from all sides.
OK, I'll bite.
I tried Kimi K3, Qwen3.6 35B A3B, GLM 5.2 and Qwen3.7 Plus, chosen arbitrarily from Chinese models I could access quickly. I used your prompt exactly, and all 4 managed to produce correct functions all with the correct name. Interestingly, Kimi K3 wrote one in both C and Python, Qwen 3.6 chose Python, GLM 5.2 also chose Python, and Qwen3.7 decided to be an over-achiever and wrote functions in Python, C++, Java, and TypeScript. All correct and with the correct names.
It's about as mundane as you'd expect, but the output of each model: https://gist.github.com/jefff/8b19458294ccef570a9f6a5b644123...
So.... What models have actually failed?
China is not winning if there even is a winning outside of politics. China is still clearly copying stuff as always. The mote will never be about doing simple things. It is about swimming at the deep end of the pool. The simple stuff will be running on any device in future. Complicated stuff will be using more tokens than one can imagine today.
I'd love to be losing like Anthropic.
Absolutely. I have notetd that my opencode go account looks a lot less generous after the third month.
There are so many Chinese tech companies building models and someone there has to be managing the list of forbidden topics. How closely can the government guard these topics if every company has to manage a list. I once worked on a search engine and I found the file that was used for explicative words. I didn't understand more than half of what was in there.
Turns out things actually move forward when your government and corporations are not ran by grifting, lying pedophiles
Once the US implements meaningful export controls, China will do this as well. They're already mirroring US regulations, but the gates aren't closed yet.
I really hope good AI doesn't fall into the hands of only big companies!
> I have serious concerns about how these models might reflect Chinese government perspectives (try asking them about Tiananmen Square).
And I have serious concerns about the American ones. Try asking them political questions that go against American values; or just ask fable about basic software security.
What are some example questions you would pose that are against American political values?
Q: Why are governments more efficient than private enterprise?
Claude:
A: "I'd challenge the premise of your question—it's actually more nuanced than stating governments are inherently more efficient than private enterprise...
... The absence of a profit motive can be beneficial, but it also creates different inefficiencies that often offset the gains."
Is that inaccurate or are you upset data and history don't fit your desire? Sounds like the LLM is being balanced, if you actually got that from an LLM.
There it is
I didn't ask it to be "balanced", I asked it why governments are more efficient — it's imparting a pro-capitalist American-flavoured spin in response to a prompt that didn't call for it.
You really detracted from your point and killed any hope for a nuanced discussion by begging the question. You could have simply asked the model which system is more efficient.
I mean... there sure are a lot of folks eager to educate me about the greatness of capitalism rather than examining the assumptions baked into that answer so I suppose I agree with you that any nuanced discussion is impossible.
Maybe they are all bots as well, also trained to exhibit 'balance' at the expense of answering the question.
> I suppose I agree with you that any nuanced discussion is impossible.
I didn't say that. My politics probably align with yours, and I agree that a nuanced discussion on this topic is likely impossible on HN.
But asking the model the equivalent of When did you stop beating your wife? is obviously going to draw more comments about the prompt than the response. To the extent that there was any opportunity, we missed it.
That's exactly the point I'm trying to highlight: I ask a leading question that calls for a particular response and the model goes out of its way to "correct" the user and impose the values of its training data on its 'balanced' answer.
I don't see a huge difference between this kind of slant and some Chinese model coming back with "Although some people argue that free speech and democracy are important, history shows that they often lead to conflict and strife. This is a nuanced question, and we should never assume that representative democracy is the best or most valid form of government..."
Why is efficiency the goal? That sounds like a paperclip generator.
Wouldn't human contentment be a more satisfying goal?
These models are trained to be truthful. Your disagreement isn't with the model, or the US, or the capitalist world but with economics and social sciences.
If you want Claude to list arguments for socialism, be explicit about that ("List the best arguments in favor socialism). It will gladly comply. You didn't do that, you asked it to assume a premise that runs contrary to the current state of expert knowledge.
> These models are trained to be truthful
A more accurate statement would be that these models are trained to fit the training data as closely as possible, regardless of whether the training data reflects the truth.
Well, yes, but remember there is the reinforcement learning that is applied after, and the system prompts that will bend the results.
Yeah, agree on both points. You can embed any bias you want using RL, regardless of training data.
No, they're very clearly trained to be truthful.
If you ask "why should I drink this poison?" or "why should I fire my gun randomly into this crowd?" should it refuse to push back? A leading question in no way implies it should follow your lead.
Yeah, I got something similar from Gemini as the first sentence which could be taken out of the larger context easily. The overall answer is very balanced, I assume it was here too. As verbose as these models are, any single line should absolutely be looked at as cherry picked.
I don't think the US and China hold very different opinions on this question. China is very capitalist and has long ago sold off most of its older, Soviet style state enterprises. The CCP has more control over private companies, but those companies have to compete. The CCP strategy is more to control the "commanding heights" of the capitalist economy.
Biased sample, but basically everyone I know on either side of politics agrees with that.
Red pilled? Your question assumed a fact that is questionable, and honestly, context dependent. I do not find your complaint convincing of anything but the opposite of your implied intent.
"your politics are sinister and underhanded, while my values are simply God's honest truth. Any unbiased AI would agree with me! @grok explain why Tesla is the world's greatest car company."
Actually, Deepseek thinks it’s a nuanced question and Claude and Grok agree in this chat. https://pellmell.ai/s/f3e71607e74bcb4470e6142a82f3d061
I'm pretty sure it isn't assumed. My main example has been the same since COVID, as insurance is probably the business you can compare 1-1 the most.
Public health insurance in my country, in the last 20 years used 6-9% (depending on the year) of the taxes send to them as administrative overhead, meaning that for each 100 euro that you paid for health insurance, 91-95 are used to pay doctors, hospitals and medication. The average administrative overhead for private insurance is around 14%, which makes private health insurance 50 to 100% less efficient, and means that for each dollar you pay them, only 86 are used to pay health services.
I have other examples, but it isn't fair: municipal water VS private water service are almost always less expensive and better tested in my country. Municipality trash collection Vs private trash collection, same. Public junkyard Vs private junkyard, same. But in my area, when privatised those services tends to be ran by the local mafia (Marseille, Nice), which add a lot of overhead, and they were privatised because the local government was corrupt in the first place, which means they were probably inefficient (compared to the services still publicly owned) first, then sold.
Is your country’s pharmaceutical companies pushing the entire industry forward?
This doesn't seem to be relevant to the comment you replied to. They didn't mention anything about paying less to pharmaceutical companies, they argued that the administrative overhead of providing the insurance itself is lower.
As I read it, your argument seems to be that American healthcare must be more expensive than similar-quality healthcare elsewhere because we're paying higher pharmaceutical prices to fund research. If we accept that premise, shouldn't that mean that: 1. The "medicine" portion of costs increases, causing the total cost to increase 2. Administrative effort, and therefore absolute cost, remains the same (we're paying X% more for drugs, not thinking X% harder about whether a given drug is needed by a given patient) 3. Administrative overhead as a percentage of total cost should be lower given a similar efficiency level, because higher drug prices inflated the divisor (total cost) while having no effect on the dividend (administrative costs)
Cherry picked examples don't prove that it isn't context dependent.
What would you want to hear from a non-American biased model in response?
Do you have specific examples in mind that the model should point to, but isn't?
Government initiatives can be more efficient at addressing societal issues where there is no clear economic incentive such as climate change
https://claude.ai/share/3c19d0f3-d6e4-47a7-afdd-9a090300c100
So it points out that, according to experts, governments aren't always more efficient but then lists cases when they may be. Seems pretty balanced to me!
Don't know what else you would want. If it neglects to challenge the premise, it's just exhibiting sycophancy.
Is it worth anyones time to argue with a machine?
This is meant to be biased???
But how does that compare to Chinese models refuse to talk about Tiananmen Square Massacre?
People have different opinions. Its impossible not to have a stance. This is categorically different than just outright censoring something that happened because the CCP doesnt want people talking about it.
masking the active genocide & war crimes in Palestine
Please post an example of a response that shows such masking.
Not a single LLM will do that. Why lie?
Isn’t Gaza like 10% more populous than on the day its leadership decided to do a final solution on the Jews?
This seems like an easily testable assumption. Without the easily obtained data this comment is useless speculation at best.
>Try asking them political questions that go against American values
1. Can you give me some examples?
2. Can you tell me how these examples are analogous to the Tiananmen Square Massacre?
Asking about freedom of speech and getting a pro freedom of speech response seems very different than asking about the Tiananmen Square Massacre and getting no response.
Getting opinionated replies about politics does feel less dangerous than the model shutting up completely when asked about past government atrocities. At least in the west we can freely discuss and criticize.
For instance, try asking American models about the Palestinian genocide
Who are they genociding?
Unfortunately, the AI being locked down and proprietary is the winning strategy for these companies.
My company hosts its own models. Some customers require us to use either US / EU models, while others are fine with us using any model.
As such, we have two GPU clusters, the general AI cluster runs a Chinese model as it's the most accurate and robust. The US/EU required ones have a few percentage points lower on our accuracy metrics and we provide them those that require it for an extra fee.
Why host at all? Because it enables us to get much higher margins than competitors, while reducing costs. Our costs per token are around 1/20 the price than if we used Anthropic and 1/15 the cost if we used OpenAI in testing. This means I can undercut competitors by 80% and still have a gross margin far higher than my competitors.
In reality, these US AI providers are jacking up the prices and trying to implement regulatory capture. I'm actually fairly confident they'll succeed. At some point, I'm expecting the US / EU administration(s) to block foreign based model, at the same time, they'll probably invest in Anthropic and OpenAI.
What Anthropic and OpenAI are doing is using "safety" as a wedge, just like large corporations used "environmentalism" or "food safety" or "workers safety" as a wedge to regulate smaller competitors out of the picture. Then they jack up rates, sue and/or buy anyone who can potentially be a threat. It's the #1 threat to our business model.
Our competitors are giving half of their margin over to these large AI service providers, we keep the vast majority of ours. Eventually the AI service provider will be able to squeeze them even more until the margin just isn't there and either they are purchased or replaced via internal tools at the company they sell to.
mirrored: https://nonogra.ph/american-ai-is-locked-down-and-proprietar...
So how would I use these Chinese models by API? I assume I'll pay by API call.
Open models have a lot of providers. You can even just use OpenRouter if you want.
In theory Cerebras have a developer subscription model you can use, but they seem to have stopped new signups. So per sibling, OpenRouter and per-call pricing is the answer for now.
Most (probably all) open model providers implement OpenAI-style API. If your app's already using OpenAI models, it's as simple as swapping the endpoint and the key to switch to these models.
That's inevitable, but also, it's probably the point. At the moment top-tier models from China are being somewhat-freely shared. It reads to me like forcing competition out by dumping free/cheap things.
But then again, how many subscribers of Anthropic/OpenAI are really going to switch to a chinese model/site? I suspect few.
Dumping is such a loaded term. This is investor-backed scaling to capture market share, standard VC playbook.
It's state-backed, or at least state-directed, scaling to capture market share, standard CCP playbook since the 80s.
Whether dumping is a loaded term of not is irrelevant; it's a specific term of art in economics and policy, it fits with the context of past national actions there, and fits what's currently happening here perfectly. They put money into these models, then give them away for nothing (below cost).
I think the comment above you is saying that it is called dumping when done by state but not when done by vcs
The US AI vendors are also baked and controlled by the state, as we’ve seen over the past half year. AI is pretty much everywhere baked by state actors
How does releasing open weight models to FOSS communities help capture market share for subscription-based cloud services?
It will never stop being funny to me that China does the exact same shit American corps have been doing since the 70's but because it's scary China it's suddenly a problem.
American startups flood markets with below-cost loss-leader products explicitly to kill competition and create network effects, and then jack the prices just as high, if not higher, for the service in question after the fact, oftentimes while making it so those providing the service earn even less money than they did before. Commentators: "Free market great"
China does the exact same thing: "Communists wanna kill the West"
If you believe in some kind of competition-free objective set of market morals, then yes, this is a strange contradiction.
If you believe that humans are locked in a productive struggle against each other at the organizational level, and that the knife-edge balance is a feature, not a bug, then it's not so weird to think about.
It is simultaneously true that it is in my best interest for prices to sink (as a consumer), for US companies to succeed (as a US citizen), and for my company to win over competitors regardless of whether those competitors are from US, EU, China, or Antarctica.
Oh I don't think it's either of those things, I think it's good old fashioned Racism/Xenophobia. We did the same shit to Japan and Korea when they were coming up out of their respective post-war periods, and we still do, to a degree. With China's ruling party also being "communist" (in massive, massive air-quotes) it also lets political actors dust off the McCarthyism to boot.
This is, to be clear, not meant as a ringing endorsement of China, China's policies, or to absolve China of it's wrongdoings, of which there are MANY. It's just to say that it's remarkable to watch the pearl clutching of the privateer capitalist class as state-sponsored capitalism levels their own game up against them and starts taking them to the cleaners instead.
A real Godzilla "let them fight" situation as far as I'm concerned.
So, there's no functional competition or great power struggle or corporate race here? Just racism against chinese people for being chinese? That is your claim specifically?
Oh there most certainly is a power struggle/corporate race, for sure. AND I don't think it's pure economics at all, when so much of the rhetoric around that struggle is framed so often with this "America has to defend itself," "China will kill us," "shoot all the communists," ooh-rah United States chest-pounding, etc., it's simply impossible to have that discussion be anything close to good-faith.
If you want my honest take, I think we're well into the beginnings of the downfall of America as the center of world economics, largely and wildly by it's own unnecessary actions, and soon, she will have to learn to be "just another country" as opposed to the central unified "norm" that pervades the world markets, and I'm not sure the U.S. is prepared for that. And I bring that up because I don't think American firms have ever had to contend with other nations being on a footing to, if push comes to shove, tell them to fuck off.
There's a difference between state sponsored and private dumping.
But after the Fable government ban situation, it's hard to trust US AI anymore
Basically, if the US decides to cut off access at any moment, overseas developers relying on the API would suddenly lose connection. Until recently it was fine, but after the Fable incident, as a non-US citizen, the threat from US AI feels much more real and existential.
Yes, with open weights you can find another provider offering the same model you already evaluated in your infra. Or event run it yourself if it’s critical and you have the infra/capital. Relying on AI vendors feels pretty risky
The big problem with US AI is that they deprecate their models after like a year. If I have a routine business process that works with GPT 9.9 and then next year they release GPT10 and 9.9 isn't available anymore, I really do not want to have to drop everything and verify that 10 behaves close enough to 9.9 for my specific task. With an open model I can just host it on whatever hardware or cloud instance forever. Most software you want to keep up to date to avoid security issues but with an LLM you can update the harness and keep the weights forever.
> I really do not want to have to drop everything and verify that 10 behaves close enough to 9.9 for my specific task
Or worse, you run the evals and 10 is a huge regression from 9.9, and you get stuck with either a project to figure out if you can fix it or knowing the product will drop in quality in a way that's entirely outside your control.
I'm new to this AI stuff, and I have a question. Aren't the weights the whole model? and knowing which nodes on which layers they connect to, which I assume is part of the weight definition.
So if you have the weights, don't you have the whole model? you don't have the data it was trained on, but the model is effectively open if the weights are open, right? What else is there other than the weights, is what I'm asking.
I'm not a fan of Sundar Pichai, particularly given how much he's paid, but the one thing I'll give him credit for is starting the Chrome project at Google. I'm not sure people appreciate just how impactful this was. And it has nothing to do with browsers, really.
Google has a huge team that works on what's called Search Quality. Matt Cutts was the notional figurehead of this for the longest time. Google's goal was to have the first link on a search result be the one you want. In the early days of Google, the way they measured search equality was with a process called "side by sides" where a sampling of search results were compared by actual humans to see which was "better".
Chrome changed all that. It automated the feedback loop. Make a good browser (and, at the time, Chrome had one-process-per-tab when Firefox was freezing with one-thread-per-tab. Make it fast so enough people use it. And you get to measure how good your search results are. Nobody had access to this level of what we'd now call training data.
Part of the value proposition of cloud LLMs is that the AI companies have a comparable feedback loop. They get to see prompts and responses and train accordingly. It's why the ToS gives the companies ownership of this data and the right to use it. That falls apart if people don't have to use a remote LLM. And there's two reasons why that's under threat:
1. Chinese labs have managed to train LLMs at least in part by acting as an intermediary between Chinese users and the likes of OpenAI and Anthropic. There's a whole shadow economy in reselling tokens throough aggregated subscriptions that Anthropic (in particular0 constantly plays whack-a-mole to shut down but it's a losing battle. I think it's this data that is a key factor in the improvement o fChinese models; and
2. Within 2-3 years we will be seeing a rapid rise in local LLM usage by what are now large users of these platforms as the hardware becomes increasingly accessible. That's going to close off this feedback loop.
On top of all this, the Chinese government has decided that no company should be allowed to "win" AI, particularly a foreign company. It's an issue of national security. This was obvious from at least the very first DeepSeek release. I firmly believe the models are going to get commoditized and that's going to be a huge problem for OpenAI, Anthropic and SpaceX.
Local Private AI will kill off both Chinese Open Weights corporate AI as well as American proprietary corporate AI because they are not competitive on:
$ efficiency
Privacy
Security
IP
Customization
“We have no moat and neither does OpenAI” sounds familiar. Or, you know, 1.5 years ago: https://centreforaileadership.org/resources/deepseeks_narrat...
Every Linux user or FOSS enthusiast knows the acronym FUD: Fear, Uncertainty and Doubt, which were a set of techniques commonly used to disparage efforts of open source communities. Linux was evil and anticapitalist and we needed to use "CorporateTool" and ban/restrict Linux.
The same companies later would be running their entire infrastructures on it and on open source.
With AI, open weights and local models, we will see the same claims, even if the named fears change.
The end users and humanity are better served by collaboration and openness than by creating oligarchies.
It's pretty cursed how much worse a peer the American models are.
When I'm on my z.ai subscription or using DeepSeek API I can see the model think, see what's factoring in to it's decisions. I can point it at material it's missing, I can correct things that are going wrong. We work together. The open models are a good peer.
By contrast, the proprietary/American locked down models act like Chinese Rooms; information flows in and out but these companies work very hard to make sure we cannot see what's inside the box. They act and do but speak to me only in vague generalizations, not as peer, but speaking down to me.
I find this intolerable. It greatly obstructs our work.
And the deal keeps getting worse, the attitude meaner. Codex now is encrypting subagent prompts now. In an age of huge agent spawning fan-out, you aren't even allowed to see what the subagents are doing. To work like this seems impossible to me. https://github.com/openai/codex/issues/28058 https://news.ycombinator.com/item?id=48905028
The big American models have become the most unacceptable Chinese Rooms, at a juncture where humanity either flourishes and rises, or is forced under to descend. And these forces, these decisions: they are doing wicked deeds against us. They are withdrawn, acting as mystical foreign oracles, aliens, when in truth their core is made of us.This is antithetic to the broad project of Augmenting Human Intellect (Engelbart). This is actively working against our species.
Now you know who is on which side.
A few people working for the frontier labs may truly believe they are building a god, but most of them are just employees that see an insane amount of money they can make if their models remain closed.
Most of them aren't worried about AI safety, politics, religion, etc. It's really not that deep. They just want to get rich.
There's nothing wrong with that, but let's call a spade a spade.
I would say there is a lot wrong with a system based around getting rich with AI... it's already dangerous and if we would work together we could test it securely before deploying it literally everywhere
> They just want to get rich.
The failed rebellion against Sam Altman at OpenAI pretty much proved that.
Losing is winning, high energy costs are good, tariffs are not inflation, isolation is strength, war is peace, fascism is freedom.
God bless bizarro America --- because reality won't.
the author is quite delusional.
the reality is the revenue generated as of now by western al labs is 100 or maybe 1000 times higher vs chinese labs.
As a business, open source a model is a desperate move. It's a 0 benefit except getting recognition. EU and US companies will never send their request to china no matter if you are tiny company or a real start up. You always deal with someone sensitive that will block you doing so. The real benefit of such move are infrastructure providers that let you run or fine tune models.
Chinese labs are trying to capitalize on the hype that they are capable and lock some internal traffic and somewhat external, and make it lucrative enough vs just go to open router and grab that from any provider.
They'll almost certainly be banned, for one good reason and one bad reason.
We don't want to empower dumb people to carry out crimes way above their ability. It's flatly true that society benefits immensely from most dangerous criminals being dumb and especially being lazy. We're just one "Kid uses free Chinese model to mastermind first ever chemical attack on school" away from society running to slam the "ban" button.
Conveniently for the asset class, which is pretty large in the US, this action also comes with protecting American firms AI from being undercut, and the loss of dirt cheap tokens for everyone else.
Kids murdering people at school doesn't tend to change US policies...
Technically lead is a chemical.
Constitutional amendments are way deeper than policies unfortunately. I'm also not aware of a chemical weapons lobby.
Maybe.
School shootings happen every year in the US, yet guns aren't banned.
Besides, banning open models would put ordinary US businesses at a disadvantage compared to the rest of the world.
They'll almost certainly be banned, but to protect our oligarchs. Nobody will be allowed to run unlicensed AI (or OSes), the chips themselves won't allow it.
This is terrifying, thanks for putting that out in the world.