
Inside a Chinese AI Lab’s Training Stack
A closer look at how leading Chinese labs are squeezing frontier-class results out of constrained hardware supply.
Wei Lian🇨🇳 China Desk LeadJun 29, 2026 6m readConstraints breed creativity, and nowhere is that clearer than in how China’s top labs approach large-scale training.
Doing more with less
Facing tighter access to the newest accelerators, several labs have leaned into aggressive quantization, custom communication kernels, and mixture-of-experts designs that activate only a fraction of parameters per token.
- Heavy use of MoE to cut active compute
- Communication-optimized training across large clusters
- Data curation treated as a first-class research problem
Why it matters globally
Many of these efficiency techniques are being published openly, and Western labs are adopting them. The constraint has, ironically, produced tooling the whole field benefits from.
Links & Resources
External links — opens in a new tab

🇨🇳 China Desk Lead · Beijing, China
Reads the Mandarin sources first — DeepSeek, Qwen, Zhipu, and the rest.

The Casio fx-CG50: A Comprehensive Academic Treatise
by Richard Murdoch Montgomery
A 223-page deep dive into hardware architecture, statistical analysis, matrix operations, and Casio BASIC programming.

Electrophysiological Biomarkers of Neuropsychiatric Brain Dynamics Vol 1
by Richard Murdoch Montgomery
EEG-based biomarkers for schizophrenia and bipolar disorder — frequency band power, event-related potentials, and neural connectivity patterns.

The TI-84 Plus C Silver Edition
by Richard Murdoch Montgomery
A 609-page volume covering arithmetic, algebra, graphing, calculus, statistics, and programming on the TI-84 Plus C Silver Edition.

A Treatise on Functional Analysis
by Richard Murdoch Montgomery
Structures, dualities, and spectra — Banach spaces, Hilbert spaces, operator theory, and spectral decompositions for the working mathematician.
Comments
Open discussion — no account needed. Be respectful.
More from Chinese Models Desk
The Claude Distillation Dispute: Seven Chinese AI Labs, Nearly 200 Million Exchanges and a Contested Model Release
Anthropic alleges that seven China-based AI labs conducted industrial-scale, unauthorized campaigns to extract capabilities from Claude, while Beijing rejects the accusations as groundless. The evidence also raises privacy, model-security and open-weights questions that the Chinese AI ecosystem will need to address.
Wei LianZhipu's GLM-5.3-FlashX Hits 200 Tokens Per Second — and the AI That Built It Is Running on Chinese Chips
Z.ai's new FlashX serving tier for GLM-5.3-Flash achieves a fourfold speed leap over the standard tier by deploying an AI agent to rewrite its own inference stack — all on a 100,000-accelerator cluster of domestically produced Chinese chips. It's the most concrete demonstration yet that China's sovereign compute ambitions are becoming real.
Sophia ChenDeepSeek-V4.1-Flash Rewrites the Rules on Agent Memory: 552B Parameters, MIT License, and a KV Cache That Fits in Your Pocket
DeepSeek's latest model slashes KV cache memory to 890 bytes per token through a radical Causal Encoder-Decoder architecture — then releases the weights for free under MIT. The result is a multimodal 552B-parameter model that outperforms its own V4-Pro flagship on agentic benchmarks at a fraction of the cost.
Wei Lian