Chinese Models Desk
Chinese Models Desk

StepFun's Step 5 Preview Bets on Million-Token Agents — and an Open-Weight Deadline

The 600-billion-parameter model combines sparse activation, multimodal input, and aggressive cache pricing to make persistent AI agents affordable. Its promised October 15 open-weight release may matter as much as its benchmark scores.

ShareWhatsAppXFacebook

# StepFun's Step 5 Preview Bets on Million-Token Agents — and an Open-Weight Deadline

StepFun, the Shanghai-based AI startup founded by former Microsoft executive Jiang Daxin, launched its most ambitious model yet on September 20, 2026. Step 5 Preview is a 600-billion-parameter sparse mixture-of-experts system designed for long-horizon agentic workflows — and it comes with a public commitment to release open weights on October 15, one week from today.

The model is available now through StepFun's hosted platform↗ and via OpenAI-compatible interfaces. But the open-weight half of the launch remains pending. Whether StepFun follows through on October 15 — and under what license — will determine whether Step 5 Preview becomes infrastructure that developers can run independently, or remains primarily a competitively priced API service.

That distinction is central to evaluating the announcement. Step 5 Preview is currently a proprietary API product with a scheduled open-weight event, not yet an open model. For now, StepFun has made a calculated bet: sparse activation, cache-aware pricing, and controllable reasoning can make persistent agents affordable enough to deploy at scale.

Sparse Scale in Service of Agent Economics

Step 5 Preview contains 600 billion total parameters but activates only 27 billion for each token — roughly 4.5% of the parameter pool participates in processing any given input. This sparse mixture-of-experts design separates representational capacity from per-token computation. Instead of applying the full model to every input, a routing mechanism selects subsets of specialized parameters, enabling economics more favorable than a dense model with a similar headline count.

The architecture is paired with a one-million-token context window and a maximum text output of 64,000 tokens. Developers can set `reasoning_effort` to `low`, `medium`, or `high`, allowing an application to trade response time and token consumption against more extensive reasoning. These controls are particularly relevant to agent systems that read large repositories, examine logs, propose patches, run tests, and iterate over many steps.

StepFun's pricing directly addresses that pattern:

  • Input (cache miss): $1.00 per million tokens — competitive with mid-tier frontier models
  • Input (cache hit): $0.05 per million tokens — one-twentieth of the cache-miss rate, making repeated context access dramatically cheaper
  • Output (including reasoning): $2.70 per million tokens — the dominant cost driver for verbose agentic workflows
  • Prompt caching: Automatically applied to stable system prompts, codebases, or document collections that persist across calls

The cache-hit rate is the most striking figure. For an agent that repeatedly reads the same large codebase or document collection, the effective input cost collapses. StepFun's thesis is clear: the relevant frontier is intelligence per completed workflow, not intelligence per isolated answer.

"The relevant frontier is intelligence per completed workflow, not intelligence per isolated answer. Step 5's cache economics are designed for agents that stay in context for hours, not seconds."

Multimodality With Concrete Limits

Step 5 Preview↗ accepts text, images, and video as inputs and produces text. The model supports as many as 60 images in one request — a material constraint for applications that analyze slide decks, interface screenshots, scanned documents, or long visual sequences.

Supported image formats include JPG/JPEG, PNG, WebP, and static GIF. StepFun recommends keeping either image dimension within 4,096 pixels. For video, the service accepts MP4, QuickTime, and Matroska files delivered through a URL, Base64 data, or StepFun's file mechanism, with files below 128MB and less than five minutes recommended.

Those boundaries mean the million-token context should not be interpreted as permission to submit arbitrarily long or large audiovisual material. Applications will still need sampling, segmentation, and retrieval strategies for longer content.

The hosted platform supports streaming, tool calling, structured JSON output, and prompt caching. Developers can access the model through Chat Completions or Messages-style interfaces, with the model ID `step-5-preview`. Compatibility reduces integration friction, but it does not make providers interchangeable — agent frameworks must still account for differences in reasoning fields, tool-call behavior, and multimodal payload limits.

The Benchmark Table Is Evidence, Not a Verdict

According to StepFun's reported evaluations↗ and third-party testing, Step 5 Preview posts strong scores across general reasoning, coding, and agent benchmarks:

  • BenchLM composite score: 70.1 out of 100, placing it fourth among Chinese models as of October 8, 2026
  • Kingbench agentic coding: 83.75% (67/80 points) in hands-on testing using the Open Code agent framework
  • DeepSWE v1.1: 67.7% — a software engineering benchmark measuring autonomous bug-fixing and feature implementation
  • Terminal-Bench 2.1: 85.0% — evaluating long-horizon terminal interaction and command-line task completion
  • ProgramBench: 80.5% — covering structured programming tasks across multiple languages
  • SciCode: 58.9% — scientific coding tasks requiring domain knowledge integration

The MindStudio coding evaluation↗ found Step 5 Preview outperforming GLM-5.3-Flash (78.75%) and MiMo-V2.6-Pro (69.38%) on Kingbench, though it trailed the full GLM-5.3 (91.25%). The model excelled at interactive tasks — creating a functional archery game with physics and leaderboards, executing a local machine learning fine-tuning pipeline — while showing relative weakness on visual 3D rendering tasks.

"Step 5 Preview's benchmark claims suggest capability, but its economics and access model are the stronger argument. On October 15, the company is due to reveal whether that argument extends beyond its own servers."

These figures are vendor-reported and have not been independently validated as a complete set. Agentic and coding results can vary significantly with scaffolding, tool permissions, retry budgets, reasoning settings, and test-time computation. For buyers, controlled tests using their own repositories, tools, latency limits, and failure costs will be more useful than an overall rank.

A Different Proposition From GLM, Kimi, and MiMo

Step 5 Preview arrives amid a rapid expansion of Chinese models built for agents and long-context work. Its closest strategic peers are not interchangeable, even where their benchmark tables overlap.

Z.ai's GLM-5.3, released August 14, 2026, emphasizes improvements achieved through post-training, with particular attention to coding, long-horizon terminal interaction, and cybersecurity tasks. Its proposition centers on specialized agent behavior and deployment through an established open-weight ecosystem. GLM-5.3 leads the Kingbench agentic coding test at 91.25%, making it the current benchmark leader for autonomous software engineering among Chinese open models.

Moonshot AI's Kimi K3 takes scale further — a 2.8-trillion-parameter multimodal MoE with a one-million-token context window, released with weights and technical documentation in July 2026. Its scale and native vision make it a broad long-horizon system, but serving such a model demands substantial multi-node infrastructure. That places practical distance between the ability to download weights and the ability to run them economically.

Xiaomi's MiMo-V2.6-Pro offers another route: a roughly one-trillion-parameter sparse MoE with multimodal support spanning text, image, video, and audio, a one-million-token context window, and an MIT-licensed open-weight release. It currently leads the BenchLM Chinese model leaderboard↗ with a score of 74.1.

Step 5 Preview's distinctive offer is the combination of 27 billion active parameters, hosted cache economics, and a promised near-term path to weights. It currently asks developers to evaluate a service; Kimi K3 and MiMo-V2.6-Pro already permit direct inspection and self-hosting. Until October 15, that access difference is more concrete than small gaps among vendor benchmark scores.

How Step 5 Compares at a Glance

  • Step 5 Preview: 600B total / 27B active params, $1.00/M input, open weights promised Oct 15, BenchLM 70.1
  • GLM-5.3 (Z.ai): 743B MoE, post-training focused, open weights available, Kingbench leader at 91.25%
  • Kimi K3 (Moonshot): 2.8T params, open weights available, BenchLM 70.5, requires multi-node serving
  • MiMo-V2.6-Pro (Xiaomi): ~1T sparse MoE, MIT license, BenchLM 74.1, broadest modality support

A Heavily Funded Lab With a Complicated Balance Sheet

StepFun was founded in April 2023↗ by Jiang Daxin, formerly Microsoft's global vice president and chief scientist at the Microsoft Software Technology Center Asia. The company's name is inspired by the mathematical "step function," representing its goal of achieving discrete leaps in AI capability. Tencent is among its investors, alongside Qiming Venture Partners, 5Y Capital, and Shanghai State-owned Capital Investment.

The company is plainly heavily funded, but available financing figures do not support a single clean cumulative total. Earlier industry reporting put cumulative funding at $1.7 billion. Later reports describe a January 2026 financing of approximately $717 million and a much larger May transaction, with some accounts citing cumulative funding above $3.2 billion. Reported valuations ranging as high as $10 billion should not be treated as settled fact on the evidence available.

The defensible conclusion is that StepFun has access to substantial capital from technology, venture, and state-linked investors. That capital matters because frontier development in China is shaped by both hardware constraints and deployment opportunity. A company must finance training and inference while also building distribution through APIs, developer tools, and devices. StepFun's low token prices can help acquire usage, but sustained economics will depend on utilization, caching behavior, infrastructure efficiency, and the reliability of the model on expensive multi-step tasks.

StepFun's Model Lineage

  • Step-2 (2024): Trillion-parameter flagship, established the company's scale ambitions
  • Step 3 (2025): Multimodal reasoning model with Multi-Matrix Factorization Attention (MFA) architecture
  • Step 3.5 Flash (2025): 196B-parameter open-source MoE, noted for efficiency and performance
  • Step 5 Preview (September 2026): Current flagship, 600B total / 27B active, multimodal, agentic focus

October 15 Is the Strategic Test

The scheduled open-weight launch carries geopolitical significance beyond StepFun's developer strategy. Chinese authorities have reportedly been considering restrictions on overseas access to the weights of advanced domestic models, with discussions including the possibility of requiring licenses for distribution or keeping frontier systems available to foreign users only through APIs.

No final policy outcome has been established. StepFun's October 15 commitment should therefore be read against an unsettled regulatory environment, not as evidence that Beijing has approved unrestricted overseas distribution or rejected future controls.

Releasing weights would create facts that are difficult to reverse once checkpoints are downloaded and mirrored. It would also allow organizations to evaluate the model independently, inspect deployment requirements, fine-tune it where permitted, and reduce dependence on StepFun's hosted service. The license will be crucial: access to files does not by itself establish broad commercial rights.

Developers interested in evaluating the model now can access it through OpenRouter↗ or directly via StepFun's API platform↗. A community-maintained BF16 checkpoint is already available on Hugging Face↗, though this is not an official StepFun release and should be treated accordingly.

If the October 15 launch proceeds with usable weights, workable inference support, and permissive terms, Step 5 Preview could become a significant option for organizations seeking million-token, multimodal agent infrastructure under their own control. If the date slips, access is geographically constrained, or the license is restrictive, the model's practical identity will remain that of an API service — capable, but not infrastructure.

The bottom line: StepFun has built a model with a credible technical proposition and unusually developer-friendly cache economics. Its benchmark claims suggest competitive capability in agentic coding and long-context reasoning. But the open-weight commitment is the real test — and that test arrives in seven days.

#StepFun#Step 5#Chinese AI#open weights#MoE#agentic AI#multimodal#benchmark#API pricing#open source
Wei Lian
Wei Lian

🇨🇳 China Desk Lead · Beijing, China

Reads the Mandarin sources first — DeepSeek, Qwen, Zhipu, and the rest.

Comments

Open discussion — no account needed. Be respectful.

0/4000
Loading comments…