StepFun released Step 5 Preview on September 20, 2026. It is a 600B parameter open model that ranks among the top three open models on the Artificial Analysis Intelligence Index, and StepFun says a single task costs about one eighth of what Claude Opus 5 charges. The full weights go public on October 15. This post covers how the model works, how it performed in StepFun’s own long running tests and in independent testing, and what to check before you plan around it.

What Step 5 Preview is under the hood

Step 5 Preview uses a sparse Mixture of Experts (MoE) architecture. It has 600B parameters in total, and only about 27B of them are activated for each token. That is roughly 4.5% of the model doing the work at any moment. The rest of the parameter pool sits idle for that token, which keeps inference costs far below what a dense 600B model would need.

Some people read “27B active” as “a 27B model.” That is a mistake. The router picks from a large pool of specialized experts, so the model’s ceiling comes from the full 600B, while the compute bill per token looks more like a mid sized model.

Other key specs:

  • Context window of 1 million tokens, large enough for an entire codebase with specs, issue history, logs and test output.
  • Native multimodal input for text and vision, including screenshots, charts and video.
  • Target workloads in AI coding, software engineering, professional knowledge work and finance.

StepFun calls the model a workhorse for real world agentic tasks. The focus is on staying coherent across long chains of tool calls, reading results, fixing errors and finishing the job, rather than answering one question well.

The Pareto frontier argument

StepFun builds its pitch around the Pareto frontier, the boundary where you can’t improve one goal without giving something up on another. For language models those goals are capability, cost and efficiency. Moving the frontier outward means getting more intelligence for the same money, or the same intelligence for less.

The company argues that scaling has mostly meant spending more compute to get more intelligence. As agent workloads grow longer and more complex, that approach gets expensive fast. StepFun’s answer is to improve how efficiently compute turns into intelligence. That idea has shaped three generations of its base models, from Step 3.5 Flash to Step 3.7 Flash and now Step 5 Preview.

Step 5 Preview pricing

The API costs $1 per million input tokens and $2.70 per million output tokens, with roughly a 95% discount on cached prompt tokens. An agent that reads 500,000 fresh input tokens and writes 50,000 output tokens would cost around $0.635 without caching. Agent loops often send the same context again and again, so caching brings the real cost of long sessions much lower.

Benchmarks and where Step 5 Preview ranks

Step 5 Preview scored 44 on the Artificial Analysis Intelligence Index, placing it in the top three open models worldwide. On the CLI subset of Agents’ Last Exam, called ALE-CLI, on the FrontierFinance investment research benchmark and on the DRACO cross domain deep research benchmark, it trails only GPT-6 Astra or Claude Opus 5 and beats every other open model tested.

The top closed models are still slightly ahead on most published benchmarks. Step 5 Preview is aimed at near frontier capability at a much lower price, not at the top spot on the leaderboard.

For coding, StepFun built StepCodeBench, which covers 553 independent repositories, 9 task types, 20 application domains and 33 programming languages. The tasks are based on industry needs, real user requests and the workflows of expert developers. Step 5 Preview stood out on overall success rate and consistency across bug fixing, feature development, refactoring and environment setup.

Long horizon agent tests

Static benchmarks can’t show whether a model keeps track of its own progress over hours of work. StepFun designed several measurable long horizon experiments to test that.

Optimizing a GPU kernel for 24 hours

Step 5 Preview got 24 hours on a single NVIDIA H100 to optimize an MLA GPU kernel from scratch. The production configuration used a head dimension of 512, batch size 1, 64 attention heads and 8,192 tokens. Starting from a written description, the model wrote changes, ran the kernel and measured throughput on its own. When a candidate ran correctly but was slower, it dropped that version and went back to its best result so far. After about 22 hours it reached 508 TFLOPS, ahead of Claude Opus 5’s 493 TFLOPS on the same task.

Automating post training

In a second 24 hour run, the model had to improve a Qwen3-30B-A3B base model on AIME24 through automated post training. It had access to an API annotator connected to production data and decided for itself how to use it, how to adjust the training data and when to iterate. Accuracy on the official AIME24 test set went from 53.3% to 60%. That matches Claude Opus 5, and Step 5 Preview used fewer annotator tokens to get there.

Hardware hacking with an ESP32

The most hands on demo turned an ESP32-S3 development board into a vibe coding keyboard with Bluetooth keys and voice input. Working from a natural language request, Step 5 Preview read large amounts of ESP32 and API documentation. It then accessed the COM port, used the camera, took device screenshots and simulated mouse input. It kept rewriting and debugging code based on what the real device reported, and ran continuously for more than three hours.

How it holds up in independent testing

A reviewer who tested Step 5 Preview before launch reports similar results, especially in frontend work. The model rebuilt an iPhone product page from a one sentence prompt in about four and a half minutes. Navigation, headlines, selling points, buttons and even the purchase links closely matched the original. The reviewer notes the model jumped 13 places from StepFun’s previous release and matched the frontend results of Kimi K3 at about a third of the price.

With reusable design Skills loaded into the agent, the model produced animated backgrounds and consistent typography and spacing in a single pass. A retro 2D platformer with heart shaped health, coins, gems and screen shake ran smoothly. A 3D version of Snake with multiple camera angles turned out just as fluid. StepFun’s own demos go further, with Blender used to build 3D assets for Three.js apps and an interactive route planner based on a 1922 Canton–Kowloon Railway travel guide.

Excel audits and long PDFs

The office tests are more relevant to most businesses. Running inside an existing Claude style agent framework, Step 5 Preview pulled figures from a 72 page PDF and calculated profit margins step by step. It extracted the text with command line tools, split it by page, built an index and logged evidence for each number.

It then audited a broken Amdocs LBO and M&A workbook with 14 sheets, more than 1,800 formulas and circular references. It first mapped what each sheet did and how the sheets fed into each other. Only then did it start looking for errors. It found 22 faulty cells in five categories and explained each one. It also listed seven cells it wasn’t confident about, marked them as uncertain and left those decisions to the user. For financial modeling, that kind of restraint is worth as much as the fixes.

Finance as a stress test

StepFun treats finance as a key proving ground because it combines macroeconomics, regulation, industry cycles and company data. The company built three internal FinStepBench evaluations for this:

  • LiveSearch tests whether the model can find and verify timely financial information as the question shifts.
  • CorporateValuation tests whether it can turn data and assumptions into consistent, reproducible valuations.
  • DeepResearch tests the full process from gathering evidence to delivering a well supported research report.

On the external FrontierFinance benchmark, which has 220 expert written questions and 11,543 grading criteria across six investment use cases, Step 5 Preview also posted strong results.

What to watch before October 15

The “Preview” label matters. Keep these points in mind:

  • License terms are still unknown. How open the weights really are depends on the rules for commercial use, modification and redistribution.
  • Self hosting needs serious hardware. Only 27B parameters are active per token, but all 600B still have to sit in memory across your hardware, so interconnect bandwidth matters.
  • Long context has a cost. Filling a million tokens increases prefill compute and latency. You still need retrieval, caching and smart context selection.
  • Tools come from the host application. The model calls tools through the framework you connect it to and has no direct access to your systems.
  • Quantization, multi node inference and fine tuning can only be properly evaluated once the weights are released.

The real test starts after the weights drop

The numbers point to a capable, inexpensive model for agent work. The detail that stands out most is the seven cells it flagged and left alone in that Excel audit. For teams handing long tasks to agents, a model that knows when to stop and ask a human may matter more than a few extra benchmark points. Once the weights are out, run it on a messy workflow of your own and check how well it reports its own uncertainty.