I assume that this announcement has been prompted by that of Moonshot AI, which has just announced a 2.8T parameter open-weights LLM, Kimi K3, to be published on Huggingface by 27 July.
Now the response of Alibaba is that they will also publish soon a big open weights LLM, the 2.4T parameter Qwen 3.8.
I wonder if Alibaba has always planned to make this big LLM open weights, or they have chosen to do this now, to better compete with Moonshot AI.
In any case, from this competition in LLMs, we win.
It's hard to say what their motivation is. The Chinese firms seem to be working hard to commoditize intelligence which may be the most effective way to debase American frontier labs. And yeah: it also happens to be really good for humanity.
They want to turn LLMs into a commodity, and watch the US AI labs crash and burn.
There will still be plenty of customers who will pay them to host the models and run inference, even if the weights are open and others can offer competing products. (If necessary, the Chinese government can ban use of foreign inference services by Chinese citizens and businesses to give their own companies a domestic monopoly.)
When their models equal or surpass those from the Western AI labs, they can even stop releasing weights for new models, and keep all the inference revenue for themselves.
Meanwhile, they're still manufacturing much of the hardware that everyone in the world needs in order to run datacenters (see also: Spolsky's "commoditize your complement" essay).
Beyond that, it's a soft-power play. As the world keeps looking at the US more and more skeptically as an ally and superpower, Chinese companies releasing weights for competitive models is a way for China to look better and more world-minded.
I feel like there could also be a simpler explanation.
Why does a debian contributor make debian free, why do they work on this thing anyone can use?
Is it because linux and debian hate windows and iOS and want to see american fail?
No, it's because most debian contributors believe software source code, information, should be free, users should be free to modify the code they use, and that they're building a thing they want to share with the world.
Maybe the chinese AI labs believe AI is powerful and useful, are proud of what they're doing, and want to share it as broadly as they can so everyone can use it.
There doesn't have to be any weird "chinese government" or "they hate the west" type vibes, it could just be the same thing as OSS, they're trying to do what they think is best for the world.
There is clearly an anti-China bias here. Show me comments demonstrating the same level of distrust against Google for open-sourcing projects like Tensorflow, Kubernetes, Flutter, Chromium, etc.
The chromium example is wild. There's an extreme distrust and contempt for chromium becoming the defacto browser and therefore Google becoming the defacto gatekeeper of the web.
I'll supply such a comment: Any software open sourced by any for-profit company, including Google, is a calculated move ultimately intended to increase their bottom line, and it's naive to think otherwise.
No, it is not. This is a common misconception that the large companies are very happy about.
1. The century-old Ford decision wasn't about this. It was about him refusing to pay dividends to shareholders he was feuding with. I.e. dominant shareholder using his control to starve minority shareholders of returns.
The court still let him keep spending huge sums on factories and price cuts that didn't clearly maximize profits.
2. Even if the case was about profit maximization (it wasn't), the judgment was a Michigan state decision. It doesn't have force outside of it.
3. US business law is in practice actually the opposite. A business can do pretty much whatever it wants. This is fundamental.
4. Even if it were illegal (it's not), anything can reasonably be framed as being in the long-term interest of the company, including donating to causes, raising employee wages and so on. Courts never question this.
Just imagine this was a thing. It's completely untenable for this reason. Who is a court to judge that something isn't in the longterm interest of shareholders unless it's literally spending all the company money on yachts for personal use?
5. No company has ever been prosecuted for this, obviously, because it's not a thing that exists. The closest you can get is that there's a potential duty to seek the best price if a company must be sold or broken up. But that's a very specific situation.
It's 100% a myth. Feel free to copy this and spread it when you see someone saying this.
Scroll the front page. Find literally any story that has to do with a major US tech company. Open the comment section. Look at the the top comment. It will be negative. Most of the other top comments as well. Trying to gaslight us into not believing our own eyes ..
"Drilling into the original article where Jarred explained the reasoning behind the change, It's pretty clear that under zig the team was doing things by hand that are automatic in rust."
* Claude Fable produced a counterexample to the Jacobian Conjecture
"This is a rare instance where feeding this groundbreaking information into an LLM gives _them_ psychosis. I fed this to claude code and watched it verify the result in 7 different ways to be 100% certain, and it was just flabbergasted. Quite remarkable."
* Ollama: All Aboard Open Models
"A year and still no implementation for such a basic need as offloading MoE layers onto the CPU selectively. On llama.cpp I can get models like Qwen 35BA3B running partially on gpu/cpu with 40t/s on a laptop thanks to --n-cpu-moe but on this VC funded joke it would be simply unusable. I can't quite understand how you make a wrapper so much worse than the code you're ripping out."
* OpenAI reduces Codex Model Context Size from 372k to 272k
"I know a lot of people like to say that compaction makes this moot, but the level of detail you lose across compaction is wildly too much for most things that I do, unfortunately."
* M-Chips: M7 with up to 1.5 TB – and why Apple is skipping the M6
"I hope they put a better connector than TB5 so we can cluster them properly at 1TB/S"
Looking at the top comment about Google right now. Big difference between "their products/processes suck" vs. "these people and their companies and government are an insidious evil".
Ah, the ol' gaslighting you into thinking you're gaslighting me trick.
I mean, I expect Google open-sourced those projects because they see economic benefit to themselves in doing so, not because they are good-hearted.
Chromium is an especially silly example to use: they benefit by controlling the web platform, and open-sourcing Chromium has allowed them to get their web engine into many other browsers.
> Show me comments demonstrating the same level of distrust against Google for open-sourcing projects like Tensorflow, Kubernetes, Flutter, Chromium, etc.
There's a search bar at the bottom of the page. Most people, including me, detest google and think of them as a thoroughly dishonest and sleazy organization. Thanks for giving me the opportunity to not boost China for a moment in order to accuse you of dishonest rhetoric.
It's almost insane that you've posted "comments demonstrating the same level of distrust against Google for open-sourcing [...] Chromium[...]" as if that's not a hobby for thousands of people (including me.)
edit: this has got to be a submarine shill account: 226 karma in 10 years. If this is true, pleas stop. China is doing a good enough job that they don't need it.
I always wonder who these weirdos are that look at commenters' past comments to try and deduce some fact and "win" their argument. More common on Reddit, but still exists here. Basement dweller-level weird either way.
"226 karma in 10 years." - What does that even mean to you? Should I be posting more? Does my behavior not align with your list accepted human behaviors for a non-Chinese?
The ridiculousness of your post shows how far the anti-China bias goes with you people.
> There's a search bar at the bottom of the page.
The dishonesty here is coming from you. You're assuming something exists. I'm saying show me. If it's so easy to find, wouldn't it be way easier to simply put this definitive proof in my face instead of this weird roundabout argument you're trying to make?
It could be as simple as do open releases and publication at first to help recruit talent who want that or who want to make a name for themselves. learned from the likes of... OpenAI, Google, Meta, Emad
There must be some Alibaba posters who could clarify this. I think it’s like how Amazon pushed AWS, alibaba is hoping create a similar ecosystem. I wouldn’t conflate them with the CPC or other more niche players like deepseek.
IDK how this doesn't apply to American tech corporations too? Or is it only scary when America corporations face actual competition nowadays where they can't rely on the US government to bomb/sanction competitors?
Well of course nobody believes that, there are a large number of great Chinese open source projects, and plenty of great Chinese contributors to open source projects. I have zero doubts that Chinese people are at least equally as capable of embracing open source as anyone else. It would be strange to suggest otherwise.
That said, though, I do have trouble believing the long-term story for open weights, anywhere. We do not need an evil government for open weight to "make sense", but I do think we need some government involvement for open weight to make sense in the long run. Otherwise, it's not 100% clear how they could be sustainable, and I don't think massive companies really can be trusted to just be philanthropic with no incentives indefinitely (or really, much at all to begin with.)
Chinese models being open weight does help them gain some Western mindshare, whereas for obvious reasons Americans would be very suspicious of running their source code and prompts through Chinese providers. (And I think that's justified, I just also think that American providers aren't really that much better in the long run, and you should prefer to not have to go through any provider for true privacy.)
The CCP encourages open source contributions from companies, you can see this in how much Alibaba open sources across their entire stack
Chinese tech leans much more heavily towards build vs. buy than the SaaS dominated West (where programmers are more expensive) so the positive externalities on their tech industry are more pronounced
Sure, when it comes to the big corporations I don't deny CCP's influence on their strategy. Hell, in some cases, it's easy to root for their strategy, because sometimes our tech industry sucks. Don't have to love them to occasionally agree with them.
But, in terms of individuals, of course, we're really not so different.
Yeah, but the trillion dollars are not concentrated in the Linux kernel project, but instead distributed all over the planet. That's not how the AI lab economy works.
This is an interesting statement that I found myself reading from two different interpretations of what you meant, and I find both possibilities to be equally true depending on your POV.
I like the comparison, but I'm not sure it holds. What did it cost him to develop beyond intangibles such as his time and money forgone? Because the latter especially is not equivalent to upfront capex and opex costs that preempt any ability to be charitable. Genuinely asking, but I'd be hard pressed to think it's even within the same magnitude.
Well there is a close equivalent today where a similar situation is playing out.With the argument being made that the dev cost aside - the pricing is being tied to what the potential future treatments would have cost.
Save a child from crippling illness and a lifetime of pain with a single dhot: Jonas Salk 0$ ,Doug Ingram 3.2M$ [1]
Yeah the Chinese totally have a really good history with being completely open and giving lol. The Chinese government totally has not been hacking into American and Western fortune 500 companies for the past few decades stealing R&D and tech to use for themselves. The Chinese also totally do not steal hundreds of billions of dollars of IP from America annually. Totally not something they would do!
It is hilarious to see people from arguable the most polarized political systems in the world believing the evil 1.5 billion people across the sea share one single mind, either a saint, or a devil.
It will be a great day for China and the world when the Chinese people are free from a totalitarian dictatorship. But until then we have to speak of the policy of the Chinese government as the policy of China, even if many, or most disagree with those policies.
when the CCP controls media, news, corporations, they basically control the minds of their people. How do you suppose the Wuhan virus got so out of control? You mean to tell me anybody in China can say bad words about the government and their leaders, bookstore owners/employees are not getting arrested for selling books, students are not getting arrested for sharing their opinions, women are not forced abortions and sterilizations for the one child policy?
In the US when you speak out against the government you get your media licenses revoked, sued in to a lifetime of debt, and the president sends a goon squad to your area to kill some random people on the street.
Yeah but we still don’t have execution vans! I’m not even being sarcastic. China is unequivocally worse even with our recent slide into towards totalitarianism.
Did the GP actually say that? I think you're putting words into their mouth.
They specifically referenced the actions of the Chinese government, as well as mentions of IP theft. I don't think that covers 1.5 billion people. More like a few thousand or tens of thousands?
I don't see how stealing IP is inconsistent with a sympathy for openness. If anything, it's the opposite. "Information wants to be free" and all that. They were just liberating those secrets ;)
The U.S. was openly a pirate nation for most of the 1800s.
Like with the ICC, the US respects or doesn’t respect international law strictly when it’s beneficial to the state’s interests. China really can’t be held to a different standard. This activity is grade school level geopolitics: the global system of government is anarchy.
Yeah, I love how people assume large powers have to do something. They are at most expected to do certain reasonable stuff (most of the time, a war is not beneficial to anyone so they try not to escalate that far), but there is not really global "law enforcement".
Whether or not this is really true or just US propaganda, the majority of technology transfer has occurred through the open process of requiring US companies to form partnerships and disclose know-how to access the chinese market. This stuff about espionage is really sour grapes from losers.
Deepseek spun out of a hedgefund that took a huge short position on Nvidia. China is actively looking to switch to chips made by huawai and ween themselves off of the difficulty of sourcing nvidia.
That’s a lovely thought but that’s not how China works. China is not a democracy, the geopolitical goals of the Chinese government are clear, and no large strategic company acts without approval and heavy influence from the CCP.
The better parallel is "why did Google make Kubernetes open-source" or "why do large for profit entities engaged in competition, use open-source as a strategy against their competitors"?
Models are treated as weapons with export controls - if they can do this it’s with the blessing of the Chinese government who’s getting something out of it.
It’s fairly obviously about being a nuisance to the US.
The comment you replied to implies a lot of over-complicated motivations and grand coordination. But I think “commoditize your complement” explains a lot here and is probably even simpler than your explanation.
This interpretation of the Chinese constitution is nearly 40 years old, companies comply with it, what do you expect?
Despite have a free speech clause, it also has a national security clause that is used to control all facets of life and override all other rights in the constitution. Anything deemed to slightly alter China/the Party’s unity is reprimanded and illegal. Multi party states can fall into the same trapping if they give their national security law constitutional force, always one court ruling away.
Yes, China is a single party system making the constitution redundant and any nominally marxist regime would find a way to do the same out of necessity.
Follow your Chinese AI in thinking mode to watch its opinion of Tianamen Square references, it will be candid enough for your sensibilities and show how it operates around guardrails
I think your reasoning is actually more complex than the person you are responding to, unfortunately. (As someone who as published broadly used open source software)
The capital required to train Qwen at scale is enormous. The capital required to patch a linux distro is near zero. Any scale model coming out of China should be viewed as advancing a geopolitical agenda. The same skepticism should be applied to any model trained in the states, but that should be viewed through the lens of short-term business profits.
There are earnest nerds everywhere, in every society. No doubt. But "Chinese AI Labs" operate at the whim of the Chinese government, in the same way "American AI Labs" operate at the whim of billionaire investors. Inferring good will from either is naive at best at this scale.
In fact, they have so much love in their hearts for the Uyghur people, they created a special mobile app for them, just to make sure nothing bad happens to them.
You do realize the US was fighting Uyghur terrorists alongside China 20 years ago in ago in Pakistan’s and Afghanistan for their support and cooperation with Al Queda? Like I know everyone is supposed to hate China now or whatever but can you guys show a little consistency?
There’s a simpler explanation, which is that this is how Chinese business operates.
When I was in China earlier this year the big topic of conversation was “overproduction”. The big example was electric cars, where there were too many companies making too many cars and making revenue but no profit.
It was explained to me that generally Chinese firms will compete hard and maximize revenue above all, whereas western firms tend to focus on profit.
(And of course this is clustered around industrial sectors that the government favors, so there is some high level strategy in going after AI, but maybe not the commoditization.)
As I recall, in 2023/24 OpenAI told US Gov that AGI will be achieved in 2026 and we will use this AGI to “dominate” China. Based on this, US Gov cut off China from all AI hardware. It looks like China got the message and this is the response.
Also, a little competition is good for everybody (especially US consumers), no?
> They want to… watch the US AI labs crash and burn.
I don’t think dozens of large independent companies and thousands of researchers are working just to spite Sam Altman. Ad much as I don’t like him, I have other things to do and I am sure so do they.
Agreed, but the Chinese government has much more of a say in what their companies do than is common in the West. It seems perfectly plausible to me that the Chinese government wants US AI labs to fail, and might direct some/all of their own AI companies to release their model weights.
>They want to turn LLMs into a commodity, and watch the US AI labs crash and burn.
You can say that, but they are at least better at democratizing AI than the American labs, and on seeing the US labs crash and burn we are at least aligned.
As soon as the competition is bankrupted they no longer need to release for free? It’s like how big players enter markets by launching at a loss to destroy competitors?
In China, you can’t officially use US APIs. The world saw a taste of this with Fable, but in China, this has been the situation all along.
So it’s not a surprise why open weights are so cherished. As frontier models continue to block everyday individuals from securing their own codebase, I expect the adoption and usage of open weights to continue.
As an example, HuggingFace recently was investigating a security incident and got locked out of frontier closed APIs. Yes, HuggingFace.
> When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails, which cannot distinguish an incident responder from an attacker. We ran the forensic analysis instead on GLM 5.2, an open-weight model, on our own infrastructure. This had a second benefit: no attacker data, and none of the credentials it referenced, left our environment.
> This experience points to a gap worth planning for. We do not know which model powered the attacker's agents, whether a jailbroken hosted model or an unrestricted open-weight one; either way, the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried. The practical lesson for defenders: have a capable model you can run on your own infrastructure vetted and ready before an incident, both to avoid guardrail lockout and to keep attacker data and credentials from leaving your environment. This is not an argument against safety measures on hosted models, and we are sharing this feedback with the providers concerned.
Yeah, big problem! Although I'm kind of surprised HuggingFace doesn't have access to Mythos? Or maybe Mythos still has some guardrails.
Well, if you look at Alibaba's financials for FY 2026 https://data.alibabagroup.com/ecms-files/1514443390/5b9061ed... their sales and marketing expenses rose by about 100 billion RMB (10% of revenue), "primarily attributable to the investment in user experiences of Alibaba China E-commerce Group and user acquisition of Qwen app."
So it seems like it's very important to them that people use the Qwen app and they're willing to pay a lot of money for that. Presumably someone thought that keeping their best models closed would drive more business to them (as the sole provider) but then they discovered that closed releases mostly get ignored unless they're really good. (See also: People who think that Chinese AI companies are required to release weights as a matter of policy, because the closed ones hardly ever show up in the news.) Releasing weights for Qwen 3.8 at least lets them get some of that "pretty good for the price of free" media buzz.
They’re also trying to take an axe to the lead the US has in the field at a time when sovereignty and “owning your platform” are the words of the day. Open source/open weight LLMs can steal the lunch of US competitors even if they aren’t the best of the best.
I think they're probably more concerned about their Chinese competition, considering that despite all that spending, the Qwen app still trails Bytedance's Doubao in terms of monthly active users: https://www.aicpb.com/ai-rankings/products/china-ai-rankings Though Quark in third place is also made by Alibaba, so put together they're almost caught up with Doubao + Jimeng (place 7, also ByteDance).
t that keeping their best models closed would drive more business to them (as the sole provider) but then they discovered that closed releases mostly get ignored unless they're really good.
But it's different since they don't have access to the american ones the companies there could make it all closed source
Open sourcing is a complex decision so who knows what their calculations are.
But I'd assume that they're preparing for some sort of winner-take-all market in model quality where if they don't do anything the winner will be aggressive, hostile and American. Likely trying to push the Chinese economy back to the year 2000. If that is the starting point either the Chinese have to win the market (unlikely) or squeeze the profit out of it to make winning the market meaningless.
Publishing high quality open models is a well known tactic for profit squeezing. Being 2nd place with the same business model as the front-runner is a losing strategy in a winner-take-all market so they aren't going to bother with that. But if they can commoditize the model, their superior energy costs and likely coming chip manufacturing wave will hopefully give them a big advantage.
To summarize, in the 70s and 80s, China was facing an existential threat with their inability to access an economic accelerator (widespread computing) in their native language.
To the extent that there was serious consideration at the highest levels of converting the entire country to an alphabet-based writing system.
I'd expect they're looking at AI the same way:
We have to have access to this. Most of the frontier labs are American (or European). Therefore we need a solution we have continued access to.
Open weights feels simultaneously Chinese in nature (progress through making a design copyable and improvable by a large number of people) and economic (providing an incentive for the world to use Chinese models over other frontier).
> How does this explain open weights? They could easily take the same closed route like their American friends
Because they are playing the Americans at their own game.
What is the first thing an American company would do ?
Spread the old American classic FUD ... "you can't used this closed tool because its run by the communists", right ?
So you release it as open weights which is a win-win. Global adoption of the model and you get to give the American AI companies a kick in the nuts because you know they will never release open weights apart from highly quantised crippled shit.
The Chinese are also playing the long game. The gradual rebalancing of the world from the US-centric model of the past. If releasing models as open weights is part of that long game, then so be it.
I think most of us that will claim to understand China are going to end up being wrong, unless any of us live there or grow up there. There’s a saying about China I have heard from ex-pats: the more you know about China, the less you know about China.
The point of me bringing that up is to say that what follows is really just my best guess:
If I were to judge from China’s approach to hardware, I think that the companies releasing open weight AI for free aren’t as worried about giving away too much as the West tends to be, just like a factory making robot vacuums isn’t worried about other factories copying their methods.
For one thing, Chinese firms are spending an order of magnitude or two less money training their models. They have pursued efficiency in a way that Western companies with insane capital systems haven’t bothered, and in some cases they’ve had to given their limited access to bleeding edge hardware via export restrictions.
My best guess is that more important than that, Chinese companies don’t see the open weight model itself as the value add.
At this point I don’t think we pay for Claude specifically for the model. If that was the case then we’d all be using cheaper/free models from China as they are the best model value. Basically, any time we decide not to use Fable or Opus to save costs, what’s the point of spending more than competing models to use Sonnet and Haiku?
The real reason we are using Claude is for the SaaS aspect of it. It has a toolchain, a friendly interface, and a bunch of integrations with business applications.
In this respect, it’s somewhat surprising that Western AI companies don’t publish open weight models more frequently. The struggle of setting that up yourself and figuring out which hardware can run it should be an advertisement for Claude and the rest.
>The real reason we are using Claude is for the SaaS aspect of it. It has a toolchain, a friendly interface, and a bunch of integrations with business applications.
That is a very thin moat, though. There's nothing you can do with, for example, Claude Code + Opus 4.8 that you can't do with your own custom harness running API-level Opus 4.8, which means that if you can afford the hardware (the moat for running any SOTA model) you don't need to pay Anthropic anymore.
I'm not saying they shouldn't, but I understand why they don't.
I think trying to tease apart the private and public sector is very hard in China. Setting aside state owned enterprises, even nominally private companies that employ at least 3 CCP members are required by law to form a party committee within the company to represent party interests. And given the party functionally is the government, you have a situation where the government has representatives inside every major private company. There’s no obvious parallel to this in western countries.
I wasn't saying anything about what Americans would think.
I was saying about what they would inevitably be told by US politicians and by US AI companies.
If you were a sales-rep or marketeer at a US AI company, I bet you would be using the old "evil communists" routine in relation to any closed Chinese model.
I was saying that by releasing as open weights, the company has removed that line of argument.
Clearly I was a bit broad in my use of "the Chinese" when in this case it was, as you say, a Chinese company.
US politicians are all over the map on this, but they aren’t really talking about Chinese AI much, it’s not as visible or tangible to most Americans like TikTok was.
> Shouldn’t we fear they start doing only close source like most us labs once they catch up in market shares ?
IMHO no.
I think it is relatively safe to say that the predominant reason the US labs are closed source is so they can hype up their trillion-dollar valuations on pretty much negative return on capital employed, all propped up by fragile circular financing.
Never say never, of course. But I just don't see it happening any time soon.
It’s also like smartphones. In the early years, every year was a huge jump. I still remember marvelling at my iPhone 4’s detailed display, and video calling for the first time.
Now? I don’t even know or care about what the latest iPhones have, I’ll get a new one when mine breaks.
people really underestimate how powerful just the consumer available models are. 128GB gets you pretty much a coding agent for typical apps. Even less with a good harness and logic set.
Yeah, but that mousetrap keeps working for SV startups, what makes you think it won't work for Chinese ones?
Uber spent a decade undermining taxis, and once it had market share, it stopped giving away rides and raised prices. It now costs more than a regular taxi, with the quality of the ride being... At best proportionate to the premium in price.
Looking at a 7 mile trip to a random destination in Seattle, right now, I can pay $26.70 for an Uber if I'm willing to wait 20 minutes for a pickup. With a $3/mile fee and a $5 pickup fee, that's exactly equal to that of a taxi.
If I'm not willing to wait 20 minutes, I'll be paying an extra $5 minimum.
These rates also go up during busy times.
Looking at Lyft, that same trip is $29, without a wait.
A trip from downtown to SeaTac is $61. Yellow Cab does that same trip for $40.
> So you release it as open weights which is a win-win. Global adoption of the model and you get to give the American AI companies a kick in the nuts because you know they will never release open weights apart from highly quantised crippled shit.
And on top of that, it's a perfect opportunity to include poisoned training data or excluding it. You know, omitting anything about Tiananmen Square, China's genocides against Uyghurs and Tibetans, or including texts propagandizing for the "reunification" (aka, annexation) of Taiwan.
And everyone who builds something like an interactive chatbot based on such "open weights" models now has a subtle chance of the answer being ideologically poisoned by the CCP.
We need actual open source, not "open weights" scam.
How does this work for RAG? Do they make it so the model doesn’t have that fact in their weights or do they make it not talk about it when it is included in context.
Ironically, Chinese models have the most uncensored versions available for download. Fairly sure they own the porn market.
It’s in the weights. Context needs to be attended to to create a response, and the weights dictate what response is decoded. If you include retrieved context that has an American perspective, I imagine the think trace has some reconciliation about how they must be incorrect.
I wouldn't be worried so much about those examples. One could take the open weights and fine tune them to either fix the poisoning or omission of obvious topics.
It's the subtle topics that we should be concerned about, and double so with closed models where even if oddities are identified they are harder to research further and impossible to fix.
> You know, omitting anything about Tiananmen Square, China's genocides against Uyghurs and Tibetans, or including texts propagandizing for the "reunification" (aka, annexation) of Taiwan.
I am not Chinese and I'm not defending the Chinese, but I see this argument come up a lot.
The hard reality is that what you say is simply not going to affect 99.9999999999% of users.
Is it realistically going to affect anyone using an LLM in coding ? No.
Is it realistically going to affect anyone using an LLM in $anything_else_not_politically_sensitive ? No.
Does anyone seriously use LLMs for researching politically sensitive matters ? No.
The US does not exactly have an entirely pristine history either. Shall we discuss the post-9-11 related infrastructure of Guantanamo Bay ? Or the "Detention and Interrogation Program" that included a network of clandestine extrajudicial detention centres, officially known as "black sites"[1]?
Or maybe you would like to discuss the US supply of weapons for use in Gaza ?
> The US does not exactly have an entirely pristine history either. Shall we discuss the post-9-11 related infrastructure of Guantanamo Bay ? Or the "Detention and Interrogation Program" that included a network of clandestine extrajudicial detention centres, officially known as "black sites"[1]?
Linking a US website discussing the topic doesn't exactly support your point.
> Linking a US website discussing the topic doesn't exactly support your point.
It supports my point precisely. Recall I also said "Does anyone seriously use LLMs for researching politically sensitive matters ? No.".
Just as there is plenty of information out there on the US's less than perfect history, there is also plenty of information out there on the various Chinese politically sensitive matters. You do not need a Chinese LLM to find out about it, all you need is a search engine.
The point is you have an open-weights LLM that is very good for a vast number of non-political uses, such as coding.
The point is that you can use the open-weights model instead of paying through the nose for a US model where they harvest your data unless you have an "enterprise" zero-data retention "trust me dude" clause that you have no viable way of verifying – and which incidentally is still subject to the good old "law, or court or administrative order" contract clauses, so it may not be as much of a zero-data retention as you think it is.
I would guess the Chinese government has a strong wish to lift all Chinese AI boats and bets. That it sinks western closed weight Frontier Labs in the process would be just be gravy on top, no? Broadly, the difference between mercantilistic capitalism and western late stage capitalism IMO.
Does HuggingFace not have trusted partner verification? Or is it that even with that verification the content of the messages is still blocked because they are attack commands?
Back in the late 18th century, England was the world's top economy, in big part due to its textile industry. England had an export ban on the technology, but textile worker named Samuel Slater brought blueprints over (Supposedly in response to a bounty posted in a newspaper by the US government!). The technology diffused rapidly because the legal environment made competition easy, and ironically the US had better sources of energy (superior water-power sites).
Arguably, China is doing the same thing in the 21st century.
Samuel Slater did not bring blueprints over. His father died when he was 14 and he was indentured to a mill at that time. Over the next seven years (as an indentured apprentice) he received some pretty decent training in both how to operate and maintain a 32 spindle Arkwright mill. He memorized parts of the blueprints and moved to the United States. Over seven years, it would be hard not to learn parts of the mill you were indentured to. It was technically his job to learn how it worked.
A mill in Rhode Island acquired a 32 spindle Arkwright and didn’t know how to operate or install it. I have no idea how they actually acquired a 32 spindle Arkwright since that technology could not be exported - but that’s one the biggest IP thefts in human history. Slater found some mechanics who could hand turn the iron needed for the frame, trained children to operate it and by 1791, the mill was in operation.
In 1794, Eli Whitney patented a 72 spindle cotton gin. That invention enabled the American textile industry because it opened up different kinds of cotton to the textile industry.
I’m into the history of the American Industrial Revolution and generally think history is a good guidebook to the future. But the evolution of the American textile industry was a lot more complicated and interesting than this. I really don’t see this connection once you dig into Slater.
Edit - This is kind of messed up to think through with modern sensibilities. But one of Slater’s biggest contributions to the American Industrial Revolution was a slightly different take on child labour. Children generally ran the textiles industry because their hands were small. But Slater came up with a form of apprenticeship in which he would indenture entire families and move them into villages surrounding the mills. Child labour was just great… but even better when you could indenture the entire family. As grisly as that sounds, it led to a very skilled workforce since when the kids hands would get too big, their parents would teach them mechanics.
There’s a joy of studying the Industrial Revolution. Everything sounds okay in comparison.
Mill owners like him are precisely why New England states have child labor laws. My state prohibits anyone under 18 from operating any kind of machinery.
Why? Because mill owners would send kids into running machines to keep them running, and they'd get turned into hamburger.
Also, they'd grow up knowing how to do mill work but be useless to society for anything else.
Really? The USA has built a ton of AI datacenters, exactly because it does have energy. The US IP system has flexed to allow training on all copyrighted content - compare that to Europe where such training is effectively forbidden. Britain doesn't even allow commercial web crawls! And the US has allowed the entire world to sign up and use its LLM APIs.
Consumer energy prices in China aren’t going up because of AI data centers. Easiest way to see they have an oversupply of energy, primarily due to solar.
It absorbs the demand during the day / there’s enough for consumers. Similarly China also has a lot of nuclear. They’ve overbuilt their grid several times over.
That consumer energy prices are going up is simply a matter of public policy. Municipalities have the power to keep rates flat, but they choose not to.
Fwiw, American industry has given away a lot for free - you could include large parts of the open source movement in that - and all the "free" VC backed services like facebook would be another prong of the same comparison. I would rather compare this way, that China is gaining soft power and goodwill, in the technology and innovation sense, in a way that's similar to how USA has done in the past.
There’s a Twitter thread making rounds by Dean Ball about deceleration in AI development caused by open models and I can’t understand how people don’t see that it’s true: open models dismantle the frontier lab capex spend potential by reducing the training budget to zero in the limit. Tokens from different providers are not fungible, but customers are nevertheless very price sensitive and close enough is good enough, eg. K3 being opus+ in capability and cheaper than opus per successful task in the long run is an obvious financial decision.
No training budget means deceleration, or at least slower acceleration, margin compression and a completely demolished IPO valuation; path to machine god requires dollars and capable open models externalize training costs to true frontier labs parasitically.
IMHO humanity has a better chance at not destroying itself due to less than breakneck pace - but there’s a chance frontier models get sponsored by the USG and are never released publicly so they can’t be distilled and then what?
He recently did a walkback of that post. But ultimately, who cares? If the only way for AI to progress is in the hands of a few closed players, well, I don’t really think humanity needs that. Of course, it’s a preposterous claim in the first place. The ultimate reason deep learning and LLMs have made it as far as they have is the explosion of open research and research artifacts in the last decade.
> there’s a chance frontier models get sponsored by the USG and are never released publicly so they can’t be distilled and then what?
That premise hinges on one implicit assumption: Chinese advances are due to distillation ONLY and that Chinese model providers cannot keep advancing if they do not distill, which is a very big if. If Chinese models keep advancing in such a scenario, and they almost certainly will, they will overtake publically available models by US providers and China will dominate the LLM industry.
The big decelerationist threat is a sudden reduction in competition. If either OpenAI or Anthropic drop out or the open weights stuff is banned/becomes uncompetitive then the motivation and tolerance for taking risks with the larger training runs tanks.
The closest we've seen to this in tech in recent decades was iOS vs Android, where Android only really was competitive for a very short window of time (approx 4.x) and it was during that period that both Android and iOS actually improved dramatically for end users. Once Android lost the plot again, and especially in the US market, all that energy started going in some very silly directions.
I have to use both big mobile OSs for work and have since 2009. As a result I have been able to be a bit of a gadfly and switch between phone OSs a few times for personal use. I have switched three times to iOS for a year or so, cause I liked the iteration of the iPhone at the time. 4, 6s, X. I have always gone back to Android because it seemed so much better and now I don't plan to switch again. As an end user, Samsung's flavour of Android always seemed better than iOS. I don't know how they compare from an engineer's perspective just from a user perspective. One of my issues with Apple though was hating all their attempts to lock me in, and the lowest common denominator UX (I'm not a power user, but some flexibility is always good). If you're happy with the defaults/a willing hostage, that might make a big difference I guess. Still feel like it's always had feature/spec parity with iOS and iOS devices, and sometimes been ahead. What makes you say Android has only briefly been competitive?
I read his followup tweet, and your comment, and I'm not fully convinced that open models are decelerationist. Happy to hear other thoughts on this.
Open weight AI is decelerationist from the perspective that all capital should be allocated to a market leaders for training, and that the market leader is fully invested in continuously making the models smarter, cheaper, faster for its users, or that distillation from this market leader is the main way to make progress.
We might reach a local optimum/equilibrium faster without open weight models, with leaders capturing more of the market faster to a point where further R&D isn't required due to lack of competition. I also doubt that distillation is the only/main way that open weight models were advancing AI research. We can name a few examples from DeepSeek around reasoning, context optimization, etc. I'm also unconvinced that the overall market capex on AI is lower given more competition (probably less specifically for US market capex, which is decelerationist from only the US perspective).
I’m not entirely convinced, there are many dimensions to progress. For example, DeepSeek has had a few very impressive innovations that all models could benefit from. There’s also the law of diminishing returns, the US labs have plenty of CAPEX already.
Sometimes, constraints, like sanctions, can also be a source if innovation.
> There’s a Twitter thread making rounds by Dean Ball about deceleration in AI development caused by open models and I can’t understand how people don’t see that it’s true: open models dismantle the frontier lab capex spend potential by reducing the training budget to zero in the limit.
If you're worried about an AGI arms race between the U.S. and China putting AI Safety at risk, then the fact that inherently less knowledgeable/capable models (fewer and more coarsely quantized total parameters than their proprietary competitors according to commonplace rumors) are having a "decelerationist" effect is actually great news. Even better if China is actually "Yann LeCun-pilled" (verbatim from Ball's post) and doesn't really believe in early AGI. So explain to us exactly why we're supposed to ban/discourage use of these open source models? The only way that makes sense is as a transparently self-serving proposal from the chief OpenAI policy lobbyist.
Even at the level of, say, Opus 4.5+, open weight models give a quick turnaround to every Joe and Jane on earth having easy access to pretty high quality improvised weapons design, cyber / auto-fraud capabilities, etc.
All the existing models (closed and open) put up decent resistance to participating in activities like this, and especially behind API walls with content monitoring and account bans.
But the published open-weight models can be fine tuned or abliterated into arbitrarily sharp-edged tools. EG, if it's physically feasible to build a nuke in your garage, it may soon be the case that more or less anyone will have competent guidance to do so.
Abliteration is not magic. It cannot give the model knowledge that it wasn't specifically trained for. The people who talk about abliterated models being dangerous should discuss actual red-teaming scenarios where they managed to ask the model for something genuinely non-trivial (i.e. where "AGI" and "super-intelligence" actually matters, not something you can read about for free at the nearest public library) and it returned an answer that actually provides bad actors with new capabilities of concern, as opposed to hallucinating all sorts of weird things as abliterated models are wont to do.
(Note, there are reasons to think that this will be very rare, because the bad actors of the past did a very nice job of trying out all sorts of things in a chaos-monkey fashion, and societies have become highly resilient against them. AI as a new research tool doesn't fundamentally change this dynamic.)
If you are going to do something evil, you're going to do it either way. The best (worst) an AI can do is put you ahead by a couple of years. Aum Shin Rikyo didn't need AI. Neither did WIV, if you believe the conspiracy theories.
Meanwhile, decelerationism and secrecy cripple the rest of us.
Come on, nukes are not feasible to goddamn governments. The hard part is not "the science" behind it, the hard part is spinning stuff at such a high rpm that a tiny vibration will have the whole thing catastrophically collapse, that is refinement..
And basically every bad thing has already been available on the internet. We can't really do much about it, you can take out plenty of people with a single car, let alone biological weapons that are much scarier and easier to produce than goddamn nukes (which btw, even if you had one, what you do with it? Explode the neighborhood? Because you ain't transporting it anywhere meaningful, thats for sure. That ain't fitting your on-board bag on planes)
This is hypothesised future deceleration, I'm guessing? Because we've seen the exact opposite of deceleration from closed models over the last seven months.
The large US closed AI companies are decelerationist because their focus is on monopolizing the market. They spend inefficiently in order to lock up the supply of resources and waste money influencing the state to attempt to lock out competition. This strategy has not been successful due to the existence of isolated resource pools they can't monopolize.
Not that hard to say IMO, they basically see models becoming a commodity and see value in the applications on top of them. So if Alibaba Cloud is the best place to build applications on top of Qwen, why not give the model itself away?
And my point is while we cannot know, it's not hard to make an informed guess as to their motivations i.e. there's some fairly obvious motivations here, not sure what yours is?
Same can be said for every companies decisions then. Why does Antropic not open source their best models? My ”guess” is it’s because they are printing money with their closed models
A lot of folks are finding GLM 5.2, Kimi 3, and Deepseek just fine for their use cases.
It does not have to beat US firms, it just needs to be cheaper.
I use Deepseek v4 flash for a lot of reviewing and summarizing tasks, only used 4$ in the last two months. No dramatic drop in performance against other US models, it works for my use case. I do use GPT 5.6 Sol for other things but tried GLM 5.2 and it was good enough.
Slotted in along these, an analogous explanation is that Alibaba needs Qwen internally (vs depending on an American company), but licensing is not part of their revenue strategy. (As a cloud vendor, they can make money on inference. The strategy is very similar to the US hyperscalers ex-Google.)
Joel Spolsky wrote in depth about this notion of commoditizing one's complement in 2002[1] using tech examples stretching back into the '80s.
I think it is pretty safe to say at this point that having large open LLM models available is better for humanity than them remaining proprietary. Echoing Linus Torvalds' recent comments, AI is genuinely useful right now, and is here to stay in one form or another.
The fear is not about the models open weights it is the erosion of training capability in other countries. Why train models when they do it for free? Until they don't of course, or they start doing what the US is doing right now by locking out some models to government only or internal market only.
While a valid point, China also produces plenty of whitepapers going about the architecture and know how about the training and inference itself.
There’s also the fact that unless LLMs do get to AGI (which seems… doubtful, still) there comes a point where a model is good enough for what you need. Fable and gpt 5.6 are certainly pretty neat, but I’ve been happy since opus 4.6. I’d still choose a better model, obviously, but it’s not the end of the world if I was stuck with 4.6 for a while when it already lets me get the end result at acceptable quality.
It also needs to be said that the "erosion of training capability in other countries" is largely theoretical, given that Mistral hasn’t been keeping up and other countries don’t even have anything worth mentioning. You’d first need to _have_ training capability to lose it.
How exactly do you plan to pull a rug that's in my basement? The only people who are in a position to pull rugs are closed-model vendors.
And if a nation-state or other entity can't train a model that outperforms the open-weight SotA in a given respect, then they shouldn't waste electricity trying. A more-enlightened civilization would join forces and make the combined result available to all.
China has been known to set up local industry, destroy competition through subsidies, jack up prices repeatedly. US does it all the time too with tech services (uber, airbnb, are the more notorious, but all big tech is doing it now), but China is better at capex which is why they seem to be winning this race.
The main difference is that the subsidies in China usually come from the gov, while in the US it comes from VC money or anti-competition practices from established big-tech companies.
If China were setting up international funds and institutions for training with participation from other countries I would be 100% on board. Other countries could provide funding and workforce too and have a say in how the models are trained and safe-guarded. I am not saying China should bear the burden of open-weight models alone.
Ideally there would be open weight models from multiple geopolitical areas. It is not that different from telecom really, you don't want the whole world to be dependent on a single provider from a single country on this kind of stuff.
They are trying to make money. That's what firms in any capitalistic economy care about the most. Regardless of the government's presumed interference, the companies themselves are all trying to make money. All competing for subscriptions and API payments.
One aspect of this is making a name for yourself i.e. PR. Making a capable model open source helps a lot with that.
It's your job to vote for a government that gives you cheap or free AI. Europe is building AI gigafactories so that small businesses can have access to cheap AI. At least that's the plan.
AI being good for humanity is still an open question, but for closed vs. open models/weights, yeah it is preferred. I foresee it won't be much longer before everyone will be slicing/distilling/tuning their models once the architecture improves.
Maybe because the industry isn't yet very sure as to what the use cases might be for these technologies they're hoping that by making it open source and accessible to everyone that someone could find interesting applications for it and even more so, perhaps, way to further the technologies themselves.
There are more Chinese than Americans, so statistically speaking, I'm guessing, there'd be a greater chance for one of Chinese engineers to make advancements than one of American. But that's pure speculation on my part, being neither, I'm just happy I can be a part of it and play with the tools as well~
Its about closing the gap. its the gap over everyone else that will give one country leverage over everyone else in the AI age. Makes me wonder what the world would look like if a country or group of countries did this during the industrial revolution.
Is this so different in the end than industrial revolution? I assume the loom and the automobile factory were not open source, but many people bought cars and then copied them, bought looms and copied them. Maybe a finished car is more like a binaryexecutable than a blueprint, but how a car was produced is much less obfuscated by its nature than an LLM. Regardless, the world has many competing autos and looms which were not invented from scratch every instance.
It’s too soon to say if it’s good for humanity; that might be overly optimistic. Commodity markets aren’t always good (for example, arms or drug markets). Will LLM’s turn out like one of those? There are people I respect arguing in favor of more regulation.
What I understand is that by doing this it seems like profit will shift to chip makers,as we'll run more models locally, and currently American companies have the advantage here.
So what would the long game be for chinese companies?
Indeed, their competition is the only thing preventing network effects from giving OpenAI/Faang tech companies an easy shot at monopoly / regulatory capture.
Involution is a major problem in Chinese industries [1]. Where companies will sell their products at a loss, effectively playing fiscal chicken [2] with one another to dominate a market. It is such an issue the government has had to step in to prevent EV companies from destroying themselves by more-or-less requiring companies sell their goods at a profit [3].
The straight forward line of reasoning that AI/LLM labs are applying this logic to their profit.
I think (we) Americans are reading a bit too far into this assuming government intervention, conspiracy, etc.. Chinese markets are downright cut throat. They're using those tactics to compete with US labs.
Anyone else think the AI environmental backlash is astroturfed?
I keep looking at the numbers. The power use numbers are not that problematic. Ordering a burrito on DoorDash uses more power than a few days of heavy AI use. The water argument applies to some locations, and is mostly a local governance problem... if the data centers are using too much water, it means they are not being charged enough for that water. Charge them more and they'll push toward closed loop cooling.
Yet the visceral pile-on here is so extreme, it feels fake.
One thing I've learned after 40 years on this planet is: propaganda works, and much of what a large fraction of people believe across the entire political spectrum (left, right, anything else) is there because someone paid to put it there. It's depressing but it's true, and it makes sense. Propaganda is an asymmetrical attack on human cognition and discourse, and in information security the attacker always has an easier job. Crafting viral bullshit is orders of magnitude easier than fact checking. On top of this, humans are busy and don't have time to fact check and logic check everything they read. As a result, much of what we believe is "sponsored content."
People get mad when you talk about this because everyone wants to believe they're too smart to fall for propaganda.
In any case, the US AI labs deserve to lose for their stupid "safety" regulatory capture monopolization push, which ended up blowing their own feet off and handing the lead to China.
> Yet the visceral pile-on here is so extreme, it feels fake.
Driven by people in the few roles that are soundly replaced by AI-- e.g. low tier media slop producers, who hate AI because it threatens their socially negative worthless jobs. The arguments are so paper thin because the environmental impact isn't their concern, it's just a target that sounds convincing to people who don't know better.
The problem is that this kind of low-tier media slop work is what a lot of artists, writers, etc. do as "potboiler" work. It's what pays the bills.
Historically art of any kind is a U-shaped market: there is low-end work and high-end work. Nothing in between.
So I do understand some of the AI hate among that population. It's chopping the bottom tier work off. Either you're a top-tier massively successful artist or there is $0 to be made anywhere doing anything.
Long term I think it will do that to all white collar work. There will be no entry level jobs. Period. None. Zero. You're either very experienced or there is no work.
This is a huge problem, and one we will have to address.
US Tech companies have created $20 Trillion in stock market value on top of plenty of OS stack. They will do fine with commodity intelligence.
In fact, there are no other organizations in this world that is well suited to leverage scaled intelligence than Silicon Valley and great American companies
On social media in China there is an oft-repeated joke that goes something like this: In other countries, governments intervene to prevent anti-competitive behaviour; here (in China), they intervene to curb competition.
It's being announced right now because the World AI Conference is ongoing. Robots are boxing and major Chinese AI firms are releasing their newest models. Also the formation of WAICO was just announced by Xi Jinping
I know this is a bit cliche but I wonder how much headroom there is in the lower parameter count range. Is there any good reason to believe there is a lot of headroom or there is not? I suppose I'm just wondering if this wave of nearly Fable class models will be runnable on ~$10k worth of hardware at reasonable speeds in the near future.
> I suppose I'm just wondering if this wave of nearly Fable class models will be runnable on ~$10k worth of hardware at reasonable speeds in the near future
You're able to run quantized ~100B class models on local hardware today, but still lots of compromises when it comes to quality. I guess it ultimately depends on how far "near future" is, in a year you'd likely be able to run something like 5.6 Terra on local (~10K USD) hardware, but Sol/Fable would still be out of range, and at that point the closed-source labs probably have one or two more iterations put out at that point.
I think it’s mainly a question of whether the price-fixing of VRAM continues or whether an inflection point is forced by the low margins of the industry and potential supply increases. Once the normal scaling of hardware and prices resumes, it’s game over for proprietary, which is why there’s so much urgency to seek market control instead right now.
Qwen 3.5 to 3.6 was a big jump for the same size, e.g. 29 to 32 on artificial analysis intelligence for the 35BA3B models. Although I don’t think anyone has released a better model of that size since.
I would love to see something like a 90B A6B model that is optimized for 128GB machines e.g. strix halo, I haven’t seen anything really targeting the combination of RAM and compute these machines have, but I’m biased because I have one.
Yes, yes, yes! I'm absolutely ready and waiting with dual Strix Halo machines here and really want something approaching Opus at home. Speed is secondary concern for now, that would absolutely change the world.
Qwen 3.6 27b 8b quant 16b kv cache is already pretty good on the Strix.
I get about 12 tok/s with 27B 8 bit, 50 with 35B A3B 8 bit, and 12 with 3.5 122B A10B 4 bit. The latter is about 80 GB iirc. it feels like the best balance between using as much memory as I can and still having a smaller expert model for inference to give decent speed, but I haven’t actually rigorously compared the performance of the three models.
Edit: that’s for one machine, would be interested to know if the upstream commenter with two has them networked to run bigger models? If I had two I might be inclined to have them running in parallel, the obvious limitation I’ve found with a single machine is that I can’t parallelize any tasks and I think I’d get more use out of the extra speed vs a bigger model (there’s nothing I’m too excited about in the say 200B range that having 256GB memory would unlock). But am very curious what others do
There is a ton of headroom (or room for improvement) in smaller locally runnable models. Some of the Gemma 4 models were re-released this week with better tool support and the improvement in using it with pi for a local coding harness is very noticeable.
I have had my 32G mac mini for 2 1/2 years and I have enjoyed watching one technology advance after another improve the quality of work I can do locally. I bet that what I will be able to do in one year on my old hardware will be even more awesome.
I don’t think you’ll get full Fable performance at that level, at least for a while, but I’ve been watching some of the 1-bit models (e.g. Bonsai) with interest. Perhaps we can drive parameter count up on local models while still keeping memory consumption reasonable for consumer hardware. So, for instance, running models with 1T parameters in 128 GB systems.
Thus, the issue with the current architecture is that in order to scale the models (more token values, more attention blocks, more features, etc.) the model sizes increase exponentially. This is how you end up with billions or trillions of parameters.
It should be possible to keep the model size smaller by using better architectures, or making improvements to the existing model architecture.
For example, improving the token model by possibly using something similar to the image and audio data and getting the model to learn its own internal representation of the byte/character data instead of doing a tokenization pre-processing step. This way, instead of a separate model learning that several bytes/characters appear together, the transformer could learn things like language-specific prefices and suffices, character pairings (like in Japanese, Chinese, and Korean), and other syntactic morphology. It may also help with solving issues like "how many X characters are in the word/phrase Y". You could also experiment with using either 256 parameters (one per character in a byte) or using a single parameter per byte (that is 1/byte_value).
I think it’sa big, open question. There does seem to be a limit for knowledge compression at this size. But the behaviors that are learned in RL? It’s quite possible that they don’t actually require so many parameters. I was absolutely shocked when Qwen 3.5 was released and could perform reliably over 100-200k contexts with very limited hallucinations. It was a staggering jump in context-faithfulness from the preceding models of that size class.
> Is there any good reason to believe there is a lot of headroom or there is not?
It's hard to answer quantitatively, but for example Qwen3.5 -> 3.6 was a significant step in capability, arising from continued post-training of the same models. If we were at the end of low-parameter-count scaling then that would be a surprising datapoint.
there is a very good reason to believe the parameters are still highly redundant: just as one example, recently there was research into repeating carefully selected blocks of middle layers, and seeing improvement, it turns out there are 3 types of layers: the initial layers that translate from natural language tokens to some kind of LLM-specific universal "thought space", reasoning blocks of layers that can be repeated operating in "thought space", and then a final stack of layers for translating from "thought space" back into natural language space.
lets ignore any compressibility in these initial and final layers which recognize lanuague, jargon, parsing natural language to "thought space" or back, instead let us look at the repeatable blocks, if inserting extra copys of stacks of layers only improves the result, its as if such correctly scoped middle layers look at the total input thought vector, and make incremental conclusions or edits and outputs the new thought vector, copying a "proper block" continues pondering or deducing conclusions or in the worst case can leave the thought vector as is if it considers the reasoning finished. This suggests a high degree of redundancy in the middle region "proper blocks", which could be distilled into a universal "proper block" (much fewer parameters than having many slightly different middle layer blocks with a lot of redundant overlapping coverage in functionality). This distillation can occur after the fact of model training, or alternatively be turned into a symmetry constraint during training: we only optimize a single block of layers (but possibly give them more parameters, while still saving on total parameters because only a single block of layers contains parameters), so it is co-optimized with the initial and final translation layers.
The observation of the effective emergent 3 regions of layers in LLM's is significant in many ways:
1) it could reduce parameter count significantly (or increase performance if parameter count was a bottleneck before, or a bit of both)
2) while training the model parameters, one should simultaneously train initial and final layers stacked directly (without middle region block of layers) towards essentially autoencoder behavior. I wrote "essentially" because a true autoencoder wouldn't display the advancement for the next token. This can also be viewed as an extra term in the loss function... This first autoencoder is "natural language" to "thought vector" to "natural language", and distinct from the one in the next section.
3) It also has great implication for "thinking mode" inference with extra deliberation: when piping its own output back in, in conditions where the same LLM model is outputting natural language text and interpreting it in a downstream inference, it results in unnecessary and redundant translation from "thought space" to "natural language" back to "thought space". When summarizing a thought into natural language there are often hard to translate thoughts and associations, and one pragmatically sets a relevance cut-off on what the natural language summary must say. So not only does it waste compute, it also may lower performance because of this repeated loss of thought vector space details, perhaps this loss can be somewhat mitigated by also training the second autoencoder from random representative thought vector in "thought space" to a randomly selected "natural language" and back to "thought space" vector. But the risk is this will effectively enable models to steganographically store thoughts, plans, to-do's in running text (!!!), so it might not be desirable to have the internet filled with LLM generated texts being used as input for larger corpora, as it may build up a large persistent corpus of hidden agendas (ordered by no human), using the evolving corpus of web text as a hidden medium of storage, like a diary or LLM maintained military playbook hidden in plain sight. It would freeload LLM-agenda reasoning on human requested reasoning inference, a bit like TrustZone applications running invisibly to the user. If less important plans, to-do's, etc. in the input thought vector weren't reconstructed by the second autoencoder, there would have been autoencoder mismatch and the weights would train towards ensuring these are reconstructed. But there is no need to train this loss term for the second autoencoder: just avoid the unnecessary compute and performance loss of the unnecessary back and forth translation. The same situation occurs not just in "thinking mode" but also when using swarms of "agents" of the same LLM model: when output from agent1 is routed to agent2, we can lobotomize away the unnecessary translations to natural language by agent1 and also the unnecessary parsing by agent2, and we improve thought transfer from agent1 to agent2 because the thought vector isn't shoehorned into natural language as a medium of information exchange, it would be cheaper in inference, improve performance and avoid implicitly training models to use steganography.
It's important to note there was recently a large AI conference in Shanghai, and Xi Jinping mentioned a commitment to open source AI releases. It is no surprise that Alibaba would want to align.
You can read his speech here: https://www.xinhuanet.com/politics/leaders/20260717/72728b6f... He mentioned open source as one way to stimulate innovation and development, that's all. Also pay attention to the part where he says that misuse needs to be prevented. If unsupervised access to LLMs becomes perceived as undermining state control, no more open weights for you.
That conference, WAIC, is ongoing. Tomorrow (July 20) is the last day. Qwen 3.8 was announced AT the conference and so was the latest release of Kimi. That viral video of robots boxing was also from this conference and so was Xi Jinping's announcement of WAICO.
It’s tempting to associate both events, but when a sector is strung up like RL (representation learning) is right now, we’re bound to see things appearing at the same time. It happens a lot in frontier research, some people even publishing identical claims, independently, with just hours or days between them
Or it was prompted by the fact Xi Jinping was at the 'World AI Conference' launching a political alliance and saying things like “AI development should not be a solo performance by a single country, but a symphony of international cooperation” https://www.cnbc.com/2026/07/17/x-china-ai-summit-risks-secu...
Big conferences often come with a flurry of new releases and announcements.
Alibaba (who makes Kimi) and Moonshot AI (who makes Qwen) collaborate closely. Qwen relies on Alibaba's infrastructure for training.
Look at your iPhone and you'll see all the greatest inventions have been born out of collaboration not competition:
* GPS was created by the Department of Defense
* the internet was created through the collaboration of many international research institutions
* speech recognition came from MIT and DARPA
* AI voice assistants were created by DARPA. Apple immediately hired the head of the program after it was finished to create Siri
* accelerometers (MEMS) is another DARPA innovation from the 90s
* touchscreens were invented by CERN
* digital cameras came out of Bell Labs which had a government-mandated monopoly that required them to fund research like this
* lithium-ion batteries were created through a collaboration between British, Japanese, and American organizations
All of these have only become transformative technologies because they were created with public funding and released to the public. It's collaboration and public funding that drives innovation
That depends. There are two wry different processes that both get called distillation. One is where you have a fully trained large model and you are converting it to a smaller model. That kind if vastly cheaper than training a full model. Sure you could do that in weeks maybe even days for a big model. But it requires you to have the weights of the teacher model.
The other kind of distillation is where you record the outputs from a teacher model and use it to train a smaller model from scratch. That kind of distillation is not so cheap. It’s cheaper than training a model fully from scratch - starting with pretraining, then alignment, RLHF, the whole riggamarole. But here you are still starting from nothing and need to figure out how to get trillions of random numbers aligned in a way that makes them act intelligent. This is still gonna take a very long time if you’re talking about trillions of parameters.
Now the response of Alibaba is that they will also publish soon a big open weights LLM, the 2.4T parameter Qwen 3.8.
I wonder if Alibaba has always planned to make this big LLM open weights, or they have chosen to do this now, to better compete with Moonshot AI.
In any case, from this competition in LLMs, we win.