ai models

China's three-week open-weight sprint: Kimi K3, Qwen3.8-Max, and the twist that could end it

In three weeks China's biggest AI labs shipped or teased frontier open-weight models — Kimi K3, Qwen3.8-Max, and DeepSeek V4 at general availability. Here is an honest read on what actually shipped, what is still a promise, and why Beijing may now restrict the wave it enabled.

T
Telli.sh Team
#open-llm#kimi-k3#qwen#deepseek#open-weight#china-ai#ai-models#meeting-notes

Between July 16 and July 26, three of China's biggest AI labs put frontier-class open models on the table. Moonshot AI announced Kimi K3 on July 16 and shipped the weights on July 26. Alibaba unveiled Qwen3.8-Max at WAIC in Shanghai on July 19. DeepSeek's V4 line reached general availability around July 20. Three labs, roughly three weeks.

The pattern is hard to miss, and it is not the one the Chinese open-model scene built its reputation on. These labs were known for small, fast, cheap models that punched above their weight. This wave is the opposite bet: trillion-parameter systems released, or promised, as open weights.

This is an honest read of that wave — what actually shipped, what is still a preview or a promise, and the twist that arrived with it: the same government that enabled these releases is now weighing whether to restrict them. If you read our July roundup, this is the sequel, and one of its cliffhangers has resolved.

Rows of cabinets inside a large supercomputer facility

Image: OLCF at ORNL, Wikimedia Commons, CC BY 2.0.

Kimi K3 is the one that actually shipped

When we wrote about Kimi K3 in July, the weights were a promise — Moonshot AI had said "by July 27." They arrived on July 26: 1.56 terabytes across 96 shards on Hugging Face under moonshotai/Kimi-K3. That is the news. An announcement is a press release; downloadable weights are a fact you can build on.

The architecture is a large mixture-of-experts model — 2.8 trillion total parameters, about 104 billion active per token, routing roughly 16 of 896 experts through a mechanism Moonshot calls Kimi Delta Attention, with a 1-million-token context window and a reported 2.5x scaling-efficiency gain over K2. On Artificial Analysis's independent ranking it placed 4th out of 189 models — behind Claude Fable 5 and two GPT-5.6 "Sol" configurations, ahead of Opus 4.8 and GPT-5.5. For an open model you can download, being in that company at all is the story.

Read the license before you plan a business around it. Kimi K3 is not MIT. It ships under a custom "Kimi K3 License" that is not OSI-approved and carries a revenue-threshold clause: run K3 as the engine of a large Model-as-a-Service business and you need a separate agreement with Moonshot. For most teams that changes nothing; for anyone building a hosted product at scale, "open weights" and "open license" turn out to be two different questions, and here only the first is an unambiguous yes.

Qwen3.8-Max is a preview, not a download

Alibaba's Qwen3.8-Max is the wave's biggest headline and its softest one. Announced at WAIC on July 19, it is a 2.4-trillion-parameter sparse MoE — active parameters undisclosed — and the first Qwen over a trillion parameters to go multimodal, taking text, image, video, and documents. It is also, right now, a hosted preview: qwen3.8-max-preview, offered at roughly 10% of standard pricing through Alibaba's Token Plan and Qoder.

Here is what has not arrived: a date for open weights, a license, a model card, or an independent benchmark table. Alibaba says the weights are coming "soon," and the Qwen Hugging Face org has had nothing new since June. The company's claim that the model is "second only to Fable 5" is exactly that — a self-report, with no third-party numbers behind it yet. Treat it as a vendor's marketing line until someone outside Alibaba measures it.

One more thing worth naming precisely, because it is already spreading as fact: the story that a "Qwen3.8 27B" drops next week is a community rumor. It traces to a Hacker News comment and one aggregator repeating it. Alibaba has named no date, no size, and no license. It is a reasonable expectation given the company's release history — and nothing more than an expectation.

DeepSeek V4 quietly reached general availability

DeepSeek did the least announcing and shipped the most usable thing. The V4 line, released back on April 24, comes in two openly licensed sizes: V4-Pro at 1.6 trillion total and 49 billion active parameters, and V4-Flash at 284 billion total and 13 billion active. Both are MIT. Both carry a 1-million-token context window and output windows up to 384K tokens. Around July 20, V4 reached general availability with time-based pricing, which is what pulled it into the same three-week frame as K3 and Qwen. There is no R2, no V5, no V4.1 — just the V4 family maturing into something teams can depend on.

The gap is closing faster than the leaderboards admit

Step back and the individual releases matter less than the slope. Nathan Lambert's read is that the distance between the best open and best closed models has compressed from roughly six-to-nine months to three-to-five — "the closest open models have been to the frontier since DeepSeek R1." The usage data points the same way. On OpenRouter, US-origin models' share of tokens fell from about 70% in June 2025 to about 30% a year later, while Chinese-origin models now regularly clear 30%. Price is a big part of why: Chinese open models tend to run 60% to 90% cheaper per token, and the hosting showed up immediately — Together AI and Modal both stood up day-zero Kimi K3 endpoints.

The quieter shift is about scale. The Chinese open-model brand was small, fast, and cheap. K3 at 2.8 trillion parameters and Qwen3.8-Max at 2.4 trillion are the exact opposite — a bet that going big in the open is now worth it. GLM-5.2 from Z.ai, released June 16 under MIT and sitting atop the open-weight tier of the Artificial Analysis Index at 51, is the same message from a fourth lab.

The twist: the government that enabled this may restrict it

Then came the turn. Around July 22-25, the Financial Times and Reuters reported that China's Ministry of Commerce is weighing export controls on AI models — including open-weight LLMs — with Alibaba, ByteDance, and Z.ai reportedly in the conversation. The framework being floated is tiered: a filing requirement for weaker models, a security review for stronger ones, and a possible ban on publicly releasing the most capable. Nothing is law yet. But the direction is striking: the wave may be cresting at the exact moment its own government starts treating these models as strategic exports rather than open contributions.

The bottom line

The tidy story is "China caught up." The more accurate one is that open weights became a strategic asset fast enough that a government now wants a hand on the tap. In three weeks, open models went from trailing the frontier by two or three quarters to trailing by one — and from being small and cheap to being some of the largest models anyone has released in any form. The open frontier did not just move. It moved somewhere its makers may not be free to keep shipping from.

For teams building on top of these models, the practical takeaway is unchanged and slightly sharper. Telli.sh already runs an open model, Qwen3, in its summarization layer, so every gain in open-model quality and price flows straight into meeting transcription, translation, and summaries. The one new line item is supply: if China's export rules land the way the reporting suggests, which open models stay freely downloadable is worth watching rather than assuming. Better models feeding the layer is the trend; where those models are allowed to come from is the thing to keep an eye on.

Start a live AI note

Sources


Back to Blog