Token Volume Is Not a Bill

Token Volume Is Not a Bill

On September 25, 2026, DeepSeek processed 54.8% of the tokens routed through Vercel AI Gateway and collected 5.4% of the spending. Anthropic processed 8.3% of the tokens and collected 38.9% of the spending.

Neither number is an estimate pulled from a chart by eye. Both come from the same machine readable file, published by the same company, for the same day, under a CC BY 4.0 licence, with no API key required. It is one of the few pieces of AI infrastructure telemetry an operator can pull, aggregate, and audit in a terminal. The interesting part is not who is winning. It is that the two boards are answering different questions, and that the answer to “who is winning” flips depending on which board an analyst happens to be reading.

The export is public, so check it yourself

The leaderboard page advertises an unauthenticated JSON export of its underlying daily series. This is what it returns for the model level dataset:

curl -s "https://vercel.com/api/ai/leaderboard-export?dataset=models&modality=all&format=json" \
  -o lb-models.json -w "http=%{http_code} bytes=%{size_download}\n"
# http=200 bytes=314087

python3 - <<'PY'
import json
d = json.load(open("lb-models.json"))
rows = d["rows"]
dates = sorted({r["date"] for r in rows})
last = dates[-1]
print("window:", dates[0], "to", last, "days:", len(dates))
for r in sorted([x for x in rows if x["date"] == last and x["metric"] == "tokens"],
                key=lambda x: -x["share_percent"])[:5]:
    print(f'{r["name"]:<22} {r["share_percent"]:.1f}%')
PY
# window: 2026-07-27 to 2026-09-25 days: 61
# DeepSeek V4.1 Flash    51.9%
# GLM 5.3 Flash          5.5%
# Gemini 3.8 Flash       3.6%
# Kimi K3                3.2%
# GPT-6 Luna             3.1%

The export carries a license field set to CC-BY-4.0, and four metrics per entity per day: tokens, requests, spend, imageCount, and videoCount. Two datasets matter here. dataset=models ranks individual models. dataset=labs rolls the same traffic up by lab, and because the shares are computed separately for each metric, the two datasets can be cross checked against each other. On the requests board for 25 September, the model level export puts Jev at 28.1% and the lab level export puts typesafe-ai at 28.3%, which is what you would expect if one model is most of its lab’s traffic. The data is internally consistent. It is the reading that goes wrong.

What each board is counting

BoardThe question it answersLeaders on 25 September 2026 (lab level)
tokensHow much work ran, counting input, output, reasoning, cached input and cache creation tokensDeepSeek 54.8%, OpenAI 10.8%, Anthropic 8.3%, Z.ai 7.6%, Google 5.8%
requestsHow many calls were madetypesafe-ai 28.3%, OpenAI 20.8%, DeepSeek 16.3%, Google 13.2%, Anthropic 6.2%
spendWhat it cost, estimated from published list pricesAnthropic 38.9%, OpenAI 24.7%, Moonshot 11.5%, Google 9.7%, Z.ai 5.9%, DeepSeek 5.4%

A token and a dollar are not the same unit, and on agent workloads the gap between them widens for structural reasons. A token count absorbs everything the loop spends: the reasoning pass, the retried tool call, the re-sent system prompt that hits cache, the fourth attempt at the same edit. A spend figure prices each of those tokens at whatever the model’s list rate happens to be. When one model costs roughly a tenth of another per million tokens, a lab can be the largest processor of work on the gateway and a rounding error in its money, at the same time, without a single figure being wrong.

Three numbers for one question

Ask Vercel’s own material what share of gateway tokens open-weight models run and you get three different answers, all defensible:

  • 56%. The September Production Index, reporting data through August 2026, states that open-weight models ran 56% of August tokens and 14% of spend, up from 7% of tokens in December 2025. It calls this the first month they took the majority of volume.
  • 70.1%. The live leaderboards page shows an open weights against closed weights split of 70.1% to 29.9% for the window labelled June 28 to September 25, 2026.
  • 78.4%. The same boards, read for the window June 21 to September 18, 2026, showed open-weight models at 78.4% of token volume against 21.6% for everything else.

Same gateway, same company, same metric name, 22 points of spread. The window explains most of it, because these are rolling views over a traffic mix that is moving quickly in both directions. The classification explains the rest, and Vercel is explicit about it in the September report’s measurement notes: open-weight classifications follow the current AI Gateway model list, which is broader than the definition used in earlier reports. Read the 78.4% as a durable fact about the market and you have taken a point in time figure, computed under one definition, on one platform, and treated it as a trend.

One more observation from the export, offered as observed rather than as a defect: the chart on the leaderboards page labels its window as starting June 28, while the unauthenticated export I pulled begins on July 27 and contains 61 daily rows. Anyone reconstructing the same window from the export instead of the chart will not reproduce the chart.

The average token price is not your bill

The same report says the average price per token across the gateway fell 23.2% in August, the third consecutive monthly drop, and that among teams running more than ten million tokens in both months the median team paid 7.6% less per token. Those two sentences sit next to each other and they are not the same claim. The 23.2% is a volume weighted average across everybody’s traffic. The 7.6% is what the middle team actually saw. A team whose token mix did not change, but whose neighbour moved a large agent workload onto a cheap model, experiences the first number as a price and the second as their own bill.

Underneath that is a measurement caveat worth quoting rather than paraphrasing. From the report’s About this report section: spending is estimated using labs’ published list prices, actual bills may differ. Every dollar figure on these boards, including the Anthropic 64% of August spend, is a list price reconstruction. It is a consistent basis for comparison and it is not an invoice. Cache discounts, committed use agreements, volume terms, and provider routing all land between the board and the bill.

The licence is not the variable

The tidy reading of the volume board against the money board is that open weights are the cheap ones, so they run the volume, while closed models sell at a premium and take the money. One entry breaks it. In the model level export for 25 September, Kimi K3 runs 3.2% of tokens and takes 10.5% of spend, and it sits second on the spend board behind Claude Opus 5.5. Kimi K3 is open-weight: Moonshot published the 2.8 trillion parameter checkpoint on Hugging Face. An open-weight model near the top of the money board cannot be explained by a licence based story.

The Kimi K3 licence is also worth reading before you treat open-weight as a synonym for free. It is a custom text, not on the Open Source Initiative approved list, and it attaches conditions at scale, including a commercial agreement above a revenue threshold. Open-weight tells you where the checkpoint lives and what you are permitted to do with it. It tells you nothing about the price per million tokens or the tier of the product line, which is what the spend board is actually ranking.

The variable that actually separates the two boards looks like tier, not licence. Models carrying Flash and Luna class labels cluster on the volume board. Models carrying Opus, Sol, and Astra class labels cluster on the money board. Open weights and low cost overlap heavily on this gateway without being the same category, and the difference matters to an operator because a licence is a procurement and deployment constraint while a tier is a routing decision.

The reading that travels furthest is the least supported

The boards do not carry an argument, but they are being used as evidence for one. An X account with a large following read the same gateway figures and concluded that open models are taking over the market beneath the frontier labs, and that this is why frontier lab founders want tighter AI regulation. In that post the account states: “This is the real reasons Dario Amodei and OpenAI want tighter AI regulation because open models are rapidly taking over the market beneath them.”

Treat that as alleged and as opinion, because that is what it is. It is one account’s interpretation, published on X, offering no source for the regulatory motive claim beyond the gateway boards themselves, and a gateway’s rolling share of its own traffic is not evidence about any lab’s policy position. It is worth naming for one reason: this is the version of the story that spreads. The measurement, with its windows and its list price caveats, stays on a JSON endpoint. The motive claim travels.

What an operator should take from this

Two metric shares from one gateway are not a market, and this one gateway is an explicit sample: developers who chose to route through it. Three consequences follow.

Route on your own numbers. The board that matters to you is your own cost per completed task, not a model’s share of somebody else’s tokens. A model at 54.8% of the world’s token volume is a statement about other teams’ workloads. If your loop is three reasoning passes and two failed tool calls per accepted change, your units of work per unit of value are your own.

Record the window and the definition beside any figure you cite. A rolling leaderboard share is a point in time value, and this one moves 22 points across three published readings. A figure presented without its window and its classification rule is not reproducible.

Do not confuse a mix shift with a discount. An average price per token falling 23.2% while the median team’s price falls 7.6% is what a compositional change looks like from the middle of the distribution. The right posture is to instrument your own spend by model and by workload, then ask whether your mix moved, rather than assuming the platform’s average came to you.

Closing

A gateway that ranks the field in one order by volume and close to the reverse order by spend is not contradicting itself. It is answering two different questions with one dataset, and both answers are correct. Token volume tells you what ran. Spend tells you, at list prices, what somebody paid. Neither one is your bill, and neither one is a market, and the operator who keeps those three sentences separate will make better routing decisions than the one who reads whichever board is at the top of the page.

Verify

Run these from any machine with curl and python3. No key, no account.

curl -s "https://vercel.com/api/ai/leaderboard-export?dataset=labs&modality=all&format=json" -o lb-labs.json
python3 - <<'PY'
import json
rows = json.load(open("lb-labs.json"))["rows"]
last = sorted({r["date"] for r in rows})[-1]
for metric in ("tokens", "requests", "spend"):
    sel = sorted([r for r in rows if r["date"] == last and r["metric"] == metric],
                 key=lambda r: -r["share_percent"])[:6]
    print(metric, last, [(r["name"], round(r["share_percent"], 1)) for r in sel])
PY

Expect tokens to lead with deepseek near 54.8%, requests to lead with typesafe-ai near 28.3%, and spend to lead with anthropic near 38.9% for 2026-09-25. If the figures have shifted, the window has moved. That is the point of the piece.

Sources

  • AI Gateway Production Index, September 2026 (data through August 2026): open-weight models at 56% of August tokens and 14% of spend, up from 7% in December 2025; average price per token down 23.2% in August with the median team down 7.6%; Anthropic at 64% of August spend; Fable 5 spend share 13.2% in July to 4.9% in August as Opus 5 rose to 22.5%; Google token share 30% to 5%; measurement notes including list price estimation and the broadened open-weight classification. Primary.
  • AI Gateway 2026 sponsor model leaderboards: open weights 70.1% against closed weights 29.9% for the window labelled June 28 to September 25, 2026; CC BY 4.0; public JSON export. Primary.
  • Public JSON export, dataset=models and dataset=labs, modality=all: every per day figure quoted for 2026-09-25, pulled in this run on 2026-09-26 and re-verified at 05:02 UTC (labs dataset, http 200, 1156178 bytes); export window 2026-07-27 to 2026-09-25, 61 daily rows. Primary.
  • FourWeekMBA, Vercel AI Gateway leaderboard: DeepSeek dominates token volume, Claude dominates spend: the June 21 to September 18, 2026 window figures (78.4% open-weight token share, DeepSeek V4.1 Flash 59.3% of tokens against 5.1% of spend, Claude Opus 4.8 at 1.7% of tokens against 13.7% of spend), and the observation that Kimi K3 publishes its weights while ranking near the top of the spend board. Secondary, citing the Vercel boards.
  • moonshotai/Kimi-K3 model card, Hugging Face: the published checkpoint, the 2.8 trillion parameter scale, and the Kimi K3 License under which the weights are released. Primary for the licence claim.
  • OpenRouter, Is Kimi K3 open source?: the licence analysis used above, that the Kimi K3 License has no SPDX identifier, is not on the OSI approved list, and attaches a model as a service revenue gate and an attribution requirement. Secondary, analysing Moonshot’s published licence.
  • X post by @MelvinInvests, 19 September 2026: the quoted regulation motive framing and the 78.4% against 21.6% reading. Attributed interpretation, unverified, presented as alleged.

Keep reading