Ranked Open Source, Filed as Proprietary

Ranked Open Source, Filed as Proprietary

On September 20, StepFun announced Step 5 Preview, a 600B-parameter sparse mixture-of-experts model with 27B parameters active per token, a 1M-token context window, and native vision input. Launch-day coverage in Sina Technology reported that the model entered the global top three among open-source models on the Artificial Analysis Intelligence Index, and that its per-task cost is one eighth that of Claude Opus 5. The article also carried the date: StepFun will open-source it on October 15.

Open the page Artificial Analysis publishes for that same model and the answer to the site’s own question is one word. “Is Step 5 Preview open source? No, Step 5 Preview is proprietary. The model weights are not publicly available.”

Nobody hid anything here. StepFun’s announcement says plainly that “the model will be released with open weights on October 15,” and Artificial Analysis classifies availability, not intentions. But the distance between a ranking and a downloadable file is exactly the distance that matters when you are deciding what to put in a routing table, and that distance is where the interesting engineering lives.

The model you can call today is not the model you can hold

Step 5 Preview is real software, and the measured numbers are good. Here is what Artificial Analysis, an independent evaluator, published for it alongside Claude Opus 5 (max effort) on the same Intelligence Index v4.3.2 run, which folds ten evaluations including Terminal-Bench 4.0, GDPval-AA v2.1, SciCode, and Humanity’s Last Exam into one score:

MetricStep 5 PreviewClaude Opus 5 (max)
Intelligence Index44 (#24 of 200)51 (#7 of 200)
Cost per index task$0.71$5.86
List price per 1M tokens$1.00 in / $2.70 out$5.00 in / $25.00 out
Output speed99.8 t/s55.3 t/s
Cache discount95%90%
Output tokens on the index160M140M
Cost to run the index evaluation$922.84$7,274.74

Divide the two cost-per-task figures and you get 8.25. StepFun’s one-eighth claim, which the company did not publish a methodology for, lands almost exactly on the ratio an outside evaluator measured. That is worth saying clearly, because the usual failure mode in vendor cost claims is that they evaporate under independent measurement. This one survives.

The caveat is the axis, not the arithmetic. The 8.25x applies across an intelligence gap of 44 against 51 index points, from the same evaluator, on the same day. StepFun’s own framing concedes the shape of it: the announcement says its task cost is lower “at a comparable level of intelligence,” while the comparison model it chose sits a tier above. And cost per task is a weighted statistic produced by running a fixed evaluation suite, not a bill. Your invoice depends on your tokens, your cache hit rate, and how much the model chooses to think. On that last point the public data has a sharp edge: Step 5 Preview generated 160M output tokens on the index against a median of 92M across comparable models, which makes it the most verbose of the three numbers in the table above. A 95% cache discount is worth a great deal if your prefixes repeat, and almost nothing if every request is a fresh long context.

The announcement’s own table argues with its headline

The composite index is where the top-three open-source framing comes from, and composites hide their components until you open them. StepFun published a full benchmark table alongside the announcement, and one row in it is worth more than the headline.

BenchmarkStep 5 Preview (High)GLM-5.3 (Max)Kimi K3 (Max)GPT-6 Astra (Max)Claude Opus 5 (Max)
Terminal-Bench v2.185.0%83.9%85.0%88.4%89.1%
Terminal-Bench v433.3%41.9%12.6%57.9%52.3%
GDPval-AA v215711634154815801735
AutomationBench-AA51.0%62.2%58.3%68.5%56.6%
DeepSWE v1.167.7%66.9%67.5%74.1%74.0%
ProgramBench (pass rate)80.5%72.0%77.8%85.4%82.3%

Two things are visible in that table. First, the same model in the same table scores 85.0% on Terminal-Bench v2.1 and 33.3% on Terminal-Bench v4. Those are two versions of one benchmark, 52 points apart, and the newer version is the one inside the composite index that produced the ranking. Benchmark version is not a footnote when the swing is larger than the gap between first and last place. Second, on that newer version the already-downloadable open-weights GLM-5.3 posts 41.9% against Step 5 Preview’s 33.3%, and the closed models sit at 52 to 58%. The composite number is defensible; the specific claim that losing the agentic row to an open competitor does not matter is the part that needs an argument, and the announcement does not make one.

The wins are real too, and they are not small. Step 5 Preview leads the table on ProgramBench at 80.5% against GLM-5.3’s 72.0%, edges DeepSWE v1.1 at 67.7% against 66.9%, and scores 88.3% on AA-LCR v1.1 in a row where GLM-5.3 manages 79.7%. The vendor also reports 24-hour long-horizon runs, a 508-TFLOPS MLA kernel optimization against Opus 5’s 493, and a Pokémon Red session sustained past 3,000 turns and six million tokens of interaction. Those last three are vendor-run experiments with no third-party reproduction yet, which does not make them false. It makes them claims.

The gap is the release pattern now, and the license rides in with the weights

StepFun is not the first lab to sell an API endpoint and the news cycle before the checkpoint. Z.ai’s GLM-5.3 launch post on August 14 promised exactly the same structure: “We will release the weights in two weeks after launch, once safety evaluation and hardening are complete.” The Hugging Face repository was created on August 25, eleven days later, and now carries 753B parameters of weights at 974,000 downloads, deployable through SGLang, vLLM, KTransformers, and Unsloth.

What landed with the weights is the part nobody could evaluate during the preview. The license file is not a standard open-source license. Hugging Face lists it as “other,” and one clause is aimed at a category most deployment teams should check themselves against: if you operate a Model as a Service business and your aggregate revenue with affiliates exceeds ten billion dollars over any consecutive twelve months, you “must pass Z.AI’s security review before using the Software or its derivative works for any commercial purpose,” with the scope and method of that review “reasonably determined by Z.AI.” Everyone below the threshold can download, modify, and self-host without approval. The point is not that the clause is unreasonable. The point is that it was knowable only after the artifact existed, while the model’s leaderboard position was broadcast a month before anyone could read the terms.

That is the honest shape of a preview release: you get the endpoint and the ranking immediately, and the deployability facts, meaning the checkpoint, its quantization, its license, and its serving story, arrive on a date the vendor chooses. StepFun’s announcement names October 15 and does not name a license.

What an operator should actually change

Nothing here argues against Step 5 Preview as an API. A 44-index model at $1.00 per million input tokens with 99.8 tokens per second of output is a genuinely competitive endpoint, and if your workloads are program synthesis or long-context retrieval the vendor data gives you reasons to test it. What the pattern argues against is treating a preview ranking as a routing option.

  • Keep a deployability column next to the capability column in your routing table, with exactly three values: downloadable now, promised with a date, or never. Step 5 Preview goes in the middle today. Nothing about its index score moves it to the first column.
  • Expire the promise. A ranking dated September 20 is a claim about September 20, and the vendor cannot guarantee the position holds by October 15, because someone else’s checkpoint can land in the meantime. Model migrations scheduled against a promised date should carry a re-evaluation trigger on that date, not a migration plan.
  • Re-price on your own traffic rather than on cost per task. Verbosity and cache hit rate dominate list price at this price tier, and a 95% cache discount is a tax on teams whose prompts do not share prefixes.
  • Run your evals against the released checkpoint, not the preview endpoint. This one is an open question rather than a finding, because neither lab has published the mapping between a preview API model and the weights that follow it, including whether the shipped build is quantized differently or post-processed before release. Until that is published, a preview evaluation is evidence about an endpoint, not about an artifact.
  • Read the license before you plan capacity for it. The terms are where these releases differ most from each other, and they are published last.

The uncomfortable truth in the spread between a chart position and a weights release is not that a vendor stretched a number. It is that an entire class of routing decisions is now made against models that do not exist as files yet, using a vocabulary (“open source”) that the evaluator running the chart explicitly refuses to apply. A ranking is a claim about relative position. An open weight is a file you can hash, quantize, audit, and serve on hardware you control. Until the file exists, the open-weight entry in your routing table is a pre-order with a date on it, and the only thing verifiable about it is the date.

Sources:

Keep reading