Chinese Models Desk
Chinese Models Desk

Free No More: Goldman Sachs Says China's AI Labs Are About to Start Charging for Their Weights

A landmark Goldman Sachs report has put a number on China's AI monetisation inflection point — and it signals that the era of unrestricted, free-for-all open weights from DeepSeek, Qwen, and GLM may be shorter than developers assumed. Here is what the shift to 'paid weights' means in practice, and which labs are moving first.

ShareWhatsAppXFacebook

Free No More: Goldman Sachs Says China's AI Labs Are About to Start Charging for Their Weights

For the past eighteen months, the defining feature of China's AI export strategy has been radical generosity. DeepSeek released its V3 and R1 weights under the MIT licence. Zhipu AI shipped GLM-5.2 — a 744-billion-parameter MoE model — under the same permissive terms. Alibaba's Qwen family has been Apache 2.0 from the start. The implicit bargain was straightforward: Chinese labs would give away the weights, capture global developer mindshare, and figure out monetisation later.

"Later" may have arrived. A landmark research report from Goldman Sachs, led by Asia internet analyst Ronald Keung, has put a precise number on the inflection point — and it is large enough to change how every developer, enterprise buyer, and cloud platform should think about their Chinese model dependencies. The bank forecasts that China's AI model API and subscription revenue will grow from approximately 35 billion RMB in 2026 to 879 billion RMB by 2030, driven by a 25-fold surge in daily token consumption. To capture that revenue, Goldman Sachs expects Chinese labs to shift from purely permissive open-source distribution toward what it calls an "open-weight plus community licence" framework — a structure that keeps weights publicly downloadable but requires commercial platforms to pay for the right to run inference on them at scale.

The report, published in July 2026, is the most detailed institutional analysis of Chinese AI monetisation to date. Its conclusions are already reshaping how enterprise procurement teams in Asia, Europe, and the Middle East are evaluating their model stacks.

---

The Numbers Behind the Shift

The strategic logic is straightforward once you see the usage data. By mid-2026, Chinese-origin models had captured a weekly peak of 46.4% of enterprise token volume on OpenRouter, the largest third-party model routing platform, up from near-zero in early 2025. When measured across only the top ten most-used models on the platform, Chinese models accounted for up to 61% of token consumption. DeepSeek alone held 17.6% of all routed tokens as of June 2026, making it the single largest vendor on the platform. Alibaba's Qwen family held 13.9%.

That adoption curve is the product of a deliberate strategy. As the US-China Commission's "Two Loops" report documented in March 2026, Chinese labs have used open-weight releases to turn the global developer community into an extension of their R&D pipeline — real-world usage generates feedback that accelerates model iteration, offsetting US advantages in raw compute. The strategy worked. The problem is that it also left enormous revenue on the table.

"Chinese AI models are achieving near-parity with US counterparts at a fraction of the cost. The question is no longer whether they can compete on performance — it is whether they can build a sustainable business model around that performance." — Goldman Sachs research note, July 2026

The pricing gap is stark. High-end Chinese models — GLM-5.2, Qwen3.7-Max, DeepSeek V4 Pro — are priced at approximately $1 per million tokens through their hosted APIs. Equivalent US frontier models from Anthropic and OpenAI run at $4–$8 per million tokens. At the low end, Chinese agent-optimised models are available for as little as $0.06–$0.20 per million tokens. That cost advantage has driven the OpenRouter share surge. But it has also meant that the labs generating the most global inference traffic are capturing only a fraction of the economic value that traffic represents.

---

What "Paid Weights" Actually Means

The Goldman Sachs report is careful to distinguish between two different things that often get conflated in discussions of open-source AI. The first is the model weights themselves — the trained parameters that define a model's capabilities. The second is the right to run inference on those weights at commercial scale.

Under the current regime, most Chinese labs release weights under licences (MIT, Apache 2.0) that permit unrestricted commercial use, including hosting the model on a cloud platform and charging customers for access. The proposed shift would not make weights private — they would remain downloadable — but would require third-party cloud platforms like AWS Bedrock, Alibaba Cloud Bailian, or Azure AI Foundry to purchase a commercial licence before offering the model as a managed inference service.

This is not a hypothetical. MiniMax has already pioneered a version of this approach, implementing revenue-sharing agreements for commercial use of its M-series models on major cloud platforms. MiniMax M3 is available on Amazon Bedrock as a fully managed model, with pricing structured through AWS's standard pay-as-you-go system — a model that allows MiniMax to capture revenue from enterprise deployments without bearing the infrastructure cost of inference itself.

Goldman Sachs expects this structure to become the industry norm. The key mechanics:

  • Weights remain publicly downloadable for individual developers, researchers, and organisations running their own infrastructure — preserving the ecosystem-building benefits of open distribution.
  • Commercial inference platforms (cloud providers, API aggregators, enterprise software vendors) would be required to purchase a licence to offer the model as a managed service, with fees structured as a revenue share or flat per-token royalty.
  • The "community licence" tier would allow non-commercial and small-scale commercial use to continue freely, maintaining developer goodwill while capturing value from high-volume enterprise deployments.
  • Frontier-tier models — the largest, most capable releases — are most likely to move to this structure first, while smaller and older models remain fully permissive.

---

Which Labs Are Moving, and How Fast

The Goldman Sachs framework evaluates Chinese AI companies across three dimensions: pricing power, cost advantages, and financial strength. Its preferred picks in the foundational text model category are Zhipu AI (the only publicly traded entity among the top tier, initiated with a target valuation of HK$1,880 per share) and DeepSeek. In multimodal and video generation, ByteDance leads, with MiniMax and Kuaishou (Kling) also rated favourably.

The readiness of each lab to execute a licensing pivot varies considerably:

  • MiniMax is the most advanced, having already established commercial licensing infrastructure through its cloud platform partnerships. Its M3 model is available on AWS Bedrock under terms that allow MiniMax to capture a share of enterprise inference revenue.
  • Zhipu AI currently distributes GLM-5.2 under a pure MIT licence with no commercial restrictions. A shift to a community licence model would require updating the terms for future releases — the existing GLM-5.2 weights, already in the wild, cannot be retroactively restricted.
  • DeepSeek has historically been the most aggressive on permissive licensing, but its pause on external fundraising and its stated focus on AGI research over commercial growth suggest the lab may be less motivated by near-term monetisation than its peers.
  • Alibaba's Qwen team has already signalled a potential strategy shift: the forthcoming Qwen3.8 — a 2.4-trillion-parameter multimodal MoE model previewed at WAIC in July — has been promised as open-weight, but Alibaba has not yet specified the licence terms. Analysts note that Alibaba's two previous flagship "Max" models remained closed, making the open-weight promise for Qwen3.8 a potential, though unverified, reversal.
"The transition from 'token maximisation' to 'ROI-focused revenue models' is the defining strategic challenge for Chinese AI labs in the second half of 2026. The labs that solve it first will have a durable competitive advantage." — Goldman Sachs, *Who Will Be the Long-Term Winner in China's AI Large Model Industry?*

---

The Regulatory Dimension

The Goldman Sachs analysis does not exist in a vacuum. It lands alongside a separate, more disruptive development: reports that China's Ministry of Commerce is consulting with major labs — including Alibaba, ByteDance, and Zhipu — on potential export controls that could restrict foreign access to frontier model weights entirely.

The two dynamics are related but distinct. The Goldman Sachs "paid weights" scenario is a commercial decision driven by monetisation logic — labs choosing to charge for what they currently give away. The MOFCOM consultation is a regulatory scenario driven by national security logic — the state potentially mandating that frontier weights not be distributed internationally at all.

Analysts expect a tiered outcome:

  • Smaller, older, and specialised models remain fully open-weight to maintain ecosystem momentum and developer goodwill.
  • Mid-tier frontier models shift to "open-weight plus community licence," capturing commercial value while preserving broad access.
  • The most capable frontier releases — models at the scale of Kimi K3 or Qwen3.8 — may face API-only distribution for international users, mirroring the approach of US frontier labs.

This tiered structure would represent a significant departure from the current landscape, where a developer in Berlin or São Paulo can download the same weights as a developer in Beijing. The CNBC analysis of how Chinese AI firms plan to monetise their free LLMs captures the tension well: the labs need revenue, but they also need the global developer ecosystem they have spent two years building.

---

What This Means for Developers and Enterprises

Practical Implications for Teams Using Chinese Models

The near-term impact on individual developers is likely to be minimal. Weights already released under MIT or Apache 2.0 cannot be retroactively restricted — GLM-5.2, DeepSeek V4, and the existing Qwen3 family will remain freely usable under their current terms. The shift affects future releases, not the current stack.

For enterprise teams, the calculus is more complex:

  • Self-hosted deployments using existing open-weight models are insulated from licensing changes. A team running GLM-5.2 on its own GPU cluster under the MIT licence has no exposure to future commercial licence requirements.
  • Managed API users on platforms like OpenRouter, AWS Bedrock, or Azure AI Foundry may see pricing adjustments as platforms pass through licence costs — though the competitive pressure from multiple Chinese labs is likely to keep price increases modest.
  • Procurement teams evaluating Chinese models for long-term enterprise contracts should now factor licence trajectory into their assessments, not just current benchmark performance and pricing.

The Benchmark Picture

Goldman Sachs' competitive framework confirms what independent leaderboards have been showing for months. Chinese models are not merely cheap alternatives — they are genuine performance competitors:

  • GLM-5.2 achieves frontier-level results on SWE-bench Pro and the Artificial Analysis Intelligence Index, rivalling Anthropic's Claude Opus 4.8 on coding benchmarks at approximately one-fifth of the cost.
  • DeepSeek V4 Pro maintains a 1-million-token context window and aggressive pricing that continues to set the global cost benchmark for long-context inference.
  • Qwen3.7-Max and the forthcoming Qwen3.8 are positioned as multimodal flagships capable of handling vision, code, and long-horizon agentic tasks within a single model family.

The efficiency underpinning these results is architectural. Goldman Sachs highlights that Chinese MoE models typically activate only 3–5% of total parameters per token — meaning a 744B-parameter model like GLM-5.2 runs inference at the effective cost of a ~30B dense model, while retaining the knowledge capacity of the full parameter count.

---

The Bigger Picture

The Goldman Sachs report is, at its core, a signal that the Chinese AI sector has crossed a maturity threshold. The "give it away and figure out monetisation later" phase is ending. What replaces it is a more nuanced commercial landscape — one where the weights may still be downloadable, but the right to build a business on top of them at scale will increasingly carry a price.

For developers, the window to lock in free, unrestricted access to frontier Chinese model weights may be narrowing. For enterprises, the time to audit which Chinese models are embedded in their stacks — and under what licence terms — is now, before the commercial landscape shifts beneath them.

The open-weight era is not over. But it is entering its second, more complicated chapter.

---

*Sources and further reading: Goldman Sachs China AI analysis via SCMP · Goldman Sachs competitive framework via HTX · OpenRouter Chinese model share data · China export controls discussion · GLM-5.2 commercial licence guide · MiniMax on AWS Bedrock*

#Goldman Sachs#China AI#Open-Weight#Commercial Licensing#DeepSeek#Qwen#Zhipu AI#MiniMax#Monetisation#Developer Tools#AI Strategy#OpenRouter
Wei Lian
Wei Lian

🇨🇳 China Desk Lead · Beijing, China

Reads the Mandarin sources first — DeepSeek, Qwen, Zhipu, and the rest.

Comments

Open discussion — no account needed. Be respectful.

0/4000
Loading comments…