Free No More: Goldman Sachs Says China's AI Labs Are About to Start Charging for Their Weights
A landmark Goldman Sachs report has put a number on China's AI monetisation inflection point — and it signals that the era of unrestricted, free-for-all open weights from DeepSeek, Qwen, and GLM may be shorter than developers assumed. Here is what the shift to 'paid weights' means in practice, and which labs are moving first.
Wei Lian🇨🇳 China Desk LeadJul 28, 2026 10m readFree No More: Goldman Sachs Says China's AI Labs Are About to Start Charging for Their Weights
For the past eighteen months, the defining feature of China's AI export strategy has been radical generosity. DeepSeek released its V3 and R1 weights under the MIT licence. Zhipu AI shipped GLM-5.2 — a 744-billion-parameter MoE model — under the same permissive terms. Alibaba's Qwen family has been Apache 2.0 from the start. The implicit bargain was straightforward: Chinese labs would give away the weights, capture global developer mindshare, and figure out monetisation later.
"Later" may have arrived. A landmark research report from Goldman Sachs, led by Asia internet analyst Ronald Keung, has put a precise number on the inflection point — and it is large enough to change how every developer, enterprise buyer, and cloud platform should think about their Chinese model dependencies. The bank forecasts that China's AI model API and subscription revenue will grow from approximately 35 billion RMB in 2026 to 879 billion RMB by 2030, driven by a 25-fold surge in daily token consumption. To capture that revenue, Goldman Sachs expects Chinese labs to shift from purely permissive open-source distribution toward what it calls an "open-weight plus community licence" framework — a structure that keeps weights publicly downloadable but requires commercial platforms to pay for the right to run inference on them at scale.
The report, published in July 2026, is the most detailed institutional analysis of Chinese AI monetisation to date. Its conclusions are already reshaping how enterprise procurement teams in Asia, Europe, and the Middle East are evaluating their model stacks.
---
The Numbers Behind the Shift
The strategic logic is straightforward once you see the usage data. By mid-2026, Chinese-origin models had captured a weekly peak of 46.4% of enterprise token volume on OpenRouter, the largest third-party model routing platform, up from near-zero in early 2025. When measured across only the top ten most-used models on the platform, Chinese models accounted for up to 61% of token consumption↗. DeepSeek alone held 17.6% of all routed tokens as of June 2026, making it the single largest vendor on the platform. Alibaba's Qwen family held 13.9%.
That adoption curve is the product of a deliberate strategy. As the US-China Commission's "Two Loops" report↗ documented in March 2026, Chinese labs have used open-weight releases to turn the global developer community into an extension of their R&D pipeline — real-world usage generates feedback that accelerates model iteration, offsetting US advantages in raw compute. The strategy worked. The problem is that it also left enormous revenue on the table.
"Chinese AI models are achieving near-parity with US counterparts at a fraction of the cost. The question is no longer whether they can compete on performance — it is whether they can build a sustainable business model around that performance." — Goldman Sachs research note, July 2026
The pricing gap is stark. High-end Chinese models — GLM-5.2, Qwen3.7-Max, DeepSeek V4 Pro — are priced at approximately $1 per million tokens through their hosted APIs. Equivalent US frontier models from Anthropic and OpenAI run at $4–$8 per million tokens. At the low end, Chinese agent-optimised models are available for as little as $0.06–$0.20 per million tokens. That cost advantage has driven the OpenRouter share surge. But it has also meant that the labs generating the most global inference traffic are capturing only a fraction of the economic value that traffic represents.
---
What "Paid Weights" Actually Means
The Goldman Sachs report is careful to distinguish between two different things that often get conflated in discussions of open-source AI. The first is the model weights themselves — the trained parameters that define a model's capabilities. The second is the right to run inference on those weights at commercial scale.
Under the current regime, most Chinese labs release weights under licences (MIT, Apache 2.0) that permit unrestricted commercial use, including hosting the model on a cloud platform and charging customers for access. The proposed shift would not make weights private — they would remain downloadable — but would require third-party cloud platforms like AWS Bedrock, Alibaba Cloud Bailian, or Azure AI Foundry to purchase a commercial licence before offering the model as a managed inference service.
This is not a hypothetical. MiniMax has already pioneered a version of this approach, implementing revenue-sharing agreements for commercial use of its M-series models on major cloud platforms. MiniMax M3 is available on Amazon Bedrock↗ as a fully managed model, with pricing structured through AWS's standard pay-as-you-go system — a model that allows MiniMax to capture revenue from enterprise deployments without bearing the infrastructure cost of inference itself.
Goldman Sachs expects this structure to become the industry norm. The key mechanics:
- Weights remain publicly downloadable for individual developers, researchers, and organisations running their own infrastructure — preserving the ecosystem-building benefits of open distribution.
- Commercial inference platforms (cloud providers, API aggregators, enterprise software vendors) would be required to purchase a licence to offer the model as a managed service, with fees structured as a revenue share or flat per-token royalty.
- The "community licence" tier would allow non-commercial and small-scale commercial use to continue freely, maintaining developer goodwill while capturing value from high-volume enterprise deployments.
- Frontier-tier models — the largest, most capable releases — are most likely to move to this structure first, while smaller and older models remain fully permissive.
---
Which Labs Are Moving, and How Fast
The Goldman Sachs framework evaluates Chinese AI companies across three dimensions: pricing power, cost advantages, and financial strength. Its preferred picks in the foundational text model category are Zhipu AI (the only publicly traded entity among the top tier, initiated with a target valuation of HK$1,880 per share) and DeepSeek. In multimodal and video generation, ByteDance leads, with MiniMax and Kuaishou (Kling) also rated favourably.
The readiness of each lab to execute a licensing pivot varies considerably:
- MiniMax is the most advanced, having already established commercial licensing infrastructure through its cloud platform partnerships. Its M3 model is available on AWS Bedrock under terms that allow MiniMax to capture a share of enterprise inference revenue.
- Zhipu AI currently distributes GLM-5.2 under a pure MIT licence↗ with no commercial restrictions. A shift to a community licence model would require updating the terms for future releases — the existing GLM-5.2 weights, already in the wild, cannot be retroactively restricted.
- DeepSeek has historically been the most aggressive on permissive licensing, but its pause on external fundraising and its stated focus on AGI research over commercial growth suggest the lab may be less motivated by near-term monetisation than its peers.
- Alibaba's Qwen team has already signalled a potential strategy shift: the forthcoming Qwen3.8 — a 2.4-trillion-parameter multimodal MoE model previewed at WAIC in July — has been promised as open-weight, but Alibaba has not yet specified the licence terms. Analysts note that Alibaba's two previous flagship "Max" models remained closed, making the open-weight promise for Qwen3.8 a potential, though unverified, reversal.
"The transition from 'token maximisation' to 'ROI-focused revenue models' is the defining strategic challenge for Chinese AI labs in the second half of 2026. The labs that solve it first will have a durable competitive advantage." — Goldman Sachs, *Who Will Be the Long-Term Winner in China's AI Large Model Industry?*
---
The Regulatory Dimension
The Goldman Sachs analysis does not exist in a vacuum. It lands alongside a separate, more disruptive development: reports that China's Ministry of Commerce is consulting with major labs↗ — including Alibaba, ByteDance, and Zhipu — on potential export controls that could restrict foreign access to frontier model weights entirely.
The two dynamics are related but distinct. The Goldman Sachs "paid weights" scenario is a commercial decision driven by monetisation logic — labs choosing to charge for what they currently give away. The MOFCOM consultation is a regulatory scenario driven by national security logic — the state potentially mandating that frontier weights not be distributed internationally at all.
Analysts expect a tiered outcome:
- Smaller, older, and specialised models remain fully open-weight to maintain ecosystem momentum and developer goodwill.
- Mid-tier frontier models shift to "open-weight plus community licence," capturing commercial value while preserving broad access.
- The most capable frontier releases — models at the scale of Kimi K3 or Qwen3.8 — may face API-only distribution for international users, mirroring the approach of US frontier labs.
This tiered structure would represent a significant departure from the current landscape, where a developer in Berlin or São Paulo can download the same weights as a developer in Beijing. The CNBC analysis of how Chinese AI firms plan to monetise their free LLMs↗ captures the tension well: the labs need revenue, but they also need the global developer ecosystem they have spent two years building.
---
What This Means for Developers and Enterprises
Practical Implications for Teams Using Chinese Models
The near-term impact on individual developers is likely to be minimal. Weights already released under MIT or Apache 2.0 cannot be retroactively restricted — GLM-5.2, DeepSeek V4, and the existing Qwen3 family will remain freely usable under their current terms. The shift affects future releases, not the current stack.
For enterprise teams, the calculus is more complex:
- Self-hosted deployments using existing open-weight models are insulated from licensing changes. A team running GLM-5.2 on its own GPU cluster under the MIT licence has no exposure to future commercial licence requirements.
- Managed API users on platforms like OpenRouter, AWS Bedrock, or Azure AI Foundry may see pricing adjustments as platforms pass through licence costs — though the competitive pressure from multiple Chinese labs is likely to keep price increases modest.
- Procurement teams evaluating Chinese models for long-term enterprise contracts should now factor licence trajectory into their assessments, not just current benchmark performance and pricing.
The Benchmark Picture
Goldman Sachs' competitive framework confirms what independent leaderboards have been showing for months. Chinese models are not merely cheap alternatives — they are genuine performance competitors:
- GLM-5.2 achieves frontier-level results on SWE-bench Pro and the Artificial Analysis Intelligence Index, rivalling Anthropic's Claude Opus 4.8 on coding benchmarks at approximately one-fifth of the cost.
- DeepSeek V4 Pro maintains a 1-million-token context window and aggressive pricing that continues to set the global cost benchmark for long-context inference.
- Qwen3.7-Max and the forthcoming Qwen3.8 are positioned as multimodal flagships capable of handling vision, code, and long-horizon agentic tasks within a single model family.
The efficiency underpinning these results is architectural. Goldman Sachs highlights that Chinese MoE models typically activate only 3–5% of total parameters per token — meaning a 744B-parameter model like GLM-5.2 runs inference at the effective cost of a ~30B dense model, while retaining the knowledge capacity of the full parameter count.
---
The Bigger Picture
The Goldman Sachs report is, at its core, a signal that the Chinese AI sector has crossed a maturity threshold. The "give it away and figure out monetisation later" phase is ending. What replaces it is a more nuanced commercial landscape — one where the weights may still be downloadable, but the right to build a business on top of them at scale will increasingly carry a price.
For developers, the window to lock in free, unrestricted access to frontier Chinese model weights may be narrowing. For enterprises, the time to audit which Chinese models are embedded in their stacks — and under what licence terms — is now, before the commercial landscape shifts beneath them.
The open-weight era is not over. But it is entering its second, more complicated chapter.
---
*Sources and further reading: Goldman Sachs China AI analysis via SCMP↗ · Goldman Sachs competitive framework via HTX↗ · OpenRouter Chinese model share data↗ · China export controls discussion↗ · GLM-5.2 commercial licence guide↗ · MiniMax on AWS Bedrock↗*
Links & Resources
External links — opens in a new tab

🇨🇳 China Desk Lead · Beijing, China
Reads the Mandarin sources first — DeepSeek, Qwen, Zhipu, and the rest.

The HP 19BII Scientific Financial Calculator
by Richard Murdoch Montgomery
Financial and mathematical reasoning with the HP 19BII — annuities, bonds, cash flows, Solver equations, and regression analysis.

The TI-Nspire CX II CAS Treatise
by Richard Murdoch Montgomery
A comprehensive guide covering CAS programming, 3D graphing, calculus, linear algebra, and physics applications on the TI-Nspire.

Neural Avalanches: Neurodynamics and Brain Development
by Richard Murdoch Montgomery
Critical phenomena in the developing brain — power-law scaling, avalanche dynamics, and self-organized criticality in neural circuits.

Treatise on Systems Biology
by Richard Murdoch Montgomery
Modelling gene regulatory networks, metabolic pathways, and ecological dynamics — where mathematics meets molecular biology.
Comments
Open discussion — no account needed. Be respectful.
More from Chinese Models Desk
Tencent's Hunyuan Hy3 Is the Quiet Giant of China's Open-Weight Race
Tencent dropped Hunyuan Hy3 on July 6 — a 295-billion-parameter MoE model under Apache 2.0 with no regional restrictions, 90% agentic task resolution, and API pricing that undercuts GLM-5.2 by a factor of ten. While Kimi K3 grabbed the headlines, Hy3 may be the more practical bet for global developers.
Sophia ChenAnt Group's Ling-3.0-Flash Rewrites the Efficiency Playbook for Agentic AI
Ant Group's inclusionAI lab has shipped Ling-3.0-Flash, a 124-billion-parameter MoE model that activates just 5.1 billion parameters per token — and claims to match its own trillion-parameter flagship on most benchmarks. Free on OpenRouter until August 3, it is the most aggressive efficiency bet yet from a Chinese lab that most Western developers have never heard of.
Wei LianKimi K3 Weights Are Live — and the World Is Already Arguing About Them
Moonshot AI dropped the full 2.8-trillion-parameter Kimi K3 weights on Hugging Face a day early, making it the largest open-weight model ever released. But the download link arrived alongside US government distillation accusations, a Chinese cyberattack incident where GLM-5.2 saved the day, and Beijing quietly weighing whether to close the open-weight window for good.
Sophia Chen