Skip to main content
MindStudio
Pricing
BlogAbout
My Workspace
Step 5 previewStepfunKingbench

Step 5 Preview Tested: Stepfun's MoE Model Scores 83.75% on Kingbench

Stepfun's Step 5 Preview MoE model takes on eight Kingbench coding tasks, from an archery game to local Gemma fine-tuning, scoring 83.75%.

Edited by Luis Chavez-Mattos, Director of Product RSS
Step 5 Preview Tested: Stepfun's MoE Model Scores 83.75% on Kingbench

What is Step 5 Preview?

Step 5 Preview is Stepfun’s new flagship large language model, built as a mixture of experts (MoE) system with 600 billion total parameters and 27 billion active per token. It supports a million tokens of context, up to 64,000 output tokens, and accepts text, image, and video input. The version currently available runs through Stepfun’s API, with full model weights planned for release on October 15. A hands-on test put it through eight coding tasks inside the Open Code agent framework, comparing its output against GLM 5.3, GLM 5.3 Flash, MIMO v2.6 Pro, and MIMO v2.6 Flash.

TL;DR

  • Step 5 Preview scored 67 out of 80 points (83.75%) across eight Kingbench coding tasks run inside Open Code, landing above GLM 5.3 Flash and both MIMO reference models but below GLM 5.3’s 91.25%.
  • The model is a 600-billion-parameter MoE with 27 billion active parameters per token, a million-token context window, and a 64,000-token output ceiling.
  • Its strongest results were a playable archery game with wind, moving targets, and a working leaderboard (9.5/10), and a local fine-tuning workflow that trained a Gemma 2B model with LoRA and served it through a local web app (9.5/10).
  • Stepfun’s API pricing runs $1 per million uncached input tokens, $5 for cached input, and $2.70 per million output tokens, with reasoning tokens billed as output.
  • Weaker spots included a 3D wristwatch that scored only 6/10 for being functional but unremarkable, and minor interaction bugs that cost points on the elevator simulation, contact lens case, and folding table demos.
  • Full model weights are scheduled for release on October 15, meaning local deployment and self-hosting become possible shortly after this API preview.

Remy doesn't build the plumbing. It inherits it.

Other agents wire up auth, databases, models, and integrations from scratch every time you ask them to build something.

200+
AI MODELS
GPT · Claude · Gemini · Llama
✓
1,000+
INTEGRATIONS
Slack · Stripe · Notion · HubSpot
✓
MANAGED DB
AUTH
PAYMENTS
CRONS

Remy ships with all of it from MindStudio — so every cycle goes into the app you actually want.

How was Step 5 Preview tested?

The test used Kingbench, a set of eight coding tasks run inside Open Code, an agentic coding environment that lets a model create files, execute commands, and inspect its own output rather than just returning a code snippet. Step 5 Preview connected directly through Stepfun’s API, while the four comparison models (GLM 5.3, GLM 5.3 Flash, MIMO v2.6 Pro, MIMO v2.6 Flash) ran through Open Router. Each task used a fresh session with the original prompt and no reference solution, high reasoning effort, and a 64,000-token output allowance.

It’s worth noting that the GLM and MIMO scores cited are historical Kingbench results, not fresh runs done in this same session. Step 5 Preview’s score comes from this specific review, which keeps the chart’s existing numbers intact while giving a direct read on how the new model performs. That distinction matters: the comparison shows where Step lands on a shared scoring scale, not a simultaneous head-to-head under identical live conditions.

What did Step 5 Preview actually build?

Across the eight tasks, Step 5 Preview produced:

  • An elevator simulation with three single-passenger cars correctly handling queued riders and random destinations (8/10, docked for overlapping floor controls and a reset bug during pickup).
  • A 3D contact lens case in Three.js with independently clickable caps and camera controls (8/10, geometry and liquid placement needed work).
  • A folding table with a slider-driven animation, though the fold direction ran backward from the prompt’s spec (8/10).
  • An SVG illustration of a panda eating a burger, clear and scalable but with some layout cleanup needed (8/10).
  • An archery game with four targets, wind effects, moving targets, a working timer, and a leaderboard that correctly sorted results and reset on replay (9.5/10).
  • A counting problem answered correctly as 20,460 (10/10).
  • A local LLM fine-tuning pipeline: Step generated a panda-facts dataset, ran a LoRA training job on a Gemma 2B Instruct model through MLX, fused the adapter, and built a local web interface that served new facts on each page refresh (9.5/10, docked for factual errors in the generated training data).
  • A 3D wristwatch with live time and two time zones, functional but visually unremarkable (6/10).

The archery game and the fine-tuning workflow stood out as the most complete demonstrations. Both required the model to coordinate multiple moving parts (physics and state for the game, data generation plus training plus inference for the fine-tuning task) and deliver something a user could directly interact with and verify.

How does Step 5 Preview compare to GLM and MIMO?

On the updated chart, Step 5 Preview’s 83.75% sits between GLM 5.3 Flash (78.75%) and GLM 5.3 (91.25%), and above both MIMO v2.6 Pro (69.38%) and MIMO v2.6 Flash (72.50%) on their original Kingbench scores. A separate MIMO Flash “math retest,” run with a larger 131,072-token output allowance, scored 80% but is tracked separately since it used a different budget than the standard runs.

In direct task comparisons noted during testing, Step 5 Preview beat a fresh GLM 5.3 Flash attempt at the archery game, where GLM’s leaderboard failed to save results properly. It also delivered a more complete local training workflow than either MIMO model’s historical fine-tuning attempt, even though the MIMO models kept their original 10/10 scores on that task in the retained chart. These are useful data points on relative capability, but they come from different test sessions and shouldn’t be read as a controlled simultaneous benchmark.

GLM 5.3 remains the top performer on this particular chart, but Step 5 Preview’s results, especially on open-ended, interactive tasks like the archery game and the fine-tuning pipeline, suggest real strength in agentic coding workflows rather than just static code generation.

Is Step 5 Preview worth trying for coding work?

For developers interested in agentic coding tools, Step 5 Preview is worth testing, particularly for tasks where you can directly verify the output: playable interactions, running simulations, or local apps backed by trained models. The fine-tuning task is a good example of why this matters. It’s not enough for a model to write training code; the pipeline has to actually run, the weights have to save correctly, and the resulting app has to load and use those weights. Step 5 Preview completed that full chain, with one correction needed for a cache setup issue in the test environment.

The pricing is also competitive for a model of this size. At $1 per million input tokens (uncached) and $2.70 per million output tokens, with a notably cheaper $5 per million cached-input rate for repeated context, it’s positioned to be usable for iterative agentic workflows where the same context gets reused across multiple calls.

The caveats are real: several demos had direction or logic bugs (the folding table moved backward from spec, the fine-tuning dataset contained factual errors), and the wristwatch task showed the model can also produce merely adequate results. Weights aren’t available for local deployment yet, scheduled for October 15, so testing today means using Stepfun’s hosted API.

Frequently Asked Questions

What is Step 5 Preview’s architecture?

It’s a mixture of experts (MoE) model with 600 billion total parameters and 27 billion active parameters per token, meaning only a fraction of the full parameter count is used for any given token during inference.

What is Kingbench?

Kingbench is a set of coding benchmark tasks (eight in this test) run inside an agentic coding environment, where a model has to create files, execute code, and produce verifiable results like playable games or running simulations, rather than just generating text.

How much does Step 5 Preview cost to use?

Stepfun’s published API pricing is $1 per million uncached input tokens, $5 per million cached input tokens, and $2.70 per million output tokens. Reasoning tokens are billed under the output rate.

When will Step 5 Preview’s weights be available?

Stepfun has stated the weights are planned for release on October 15, after which local deployment should become possible alongside the existing API access.

How does Step 5 Preview compare to GLM 5.3?

In this test, Step 5 Preview scored 83.75% against GLM 5.3’s 91.25% on the same Kingbench scoring scale, with GLM 5.3’s figure drawn from historical results rather than a fresh simultaneous run. GLM 5.3 remains ahead overall, though Step 5 Preview outperformed GLM 5.3 Flash and both MIMO models on this chart.

Editorial standards

Presented by MindStudio

No spam. Unsubscribe anytime.