The Two Dollar Model Is Not the Price Floor

The Two Dollar Model Is Not the Price Floor

A chart circulating today puts three models at $2 per million input tokens: GPT-6 Sol, Claude Sonnet 5 and Grok 4.7. It calls that the settled mid-tier, and GPT-6 Luna the ten-cent floor.

The prices are a snapshot, not the whole market. Here is the fuller comparison, including Chinese APIs and the conditions that change the sticker price.

Price comparison

Prices are per million tokens. Dollar and yuan prices are kept separate. The rates below are published list prices, not a prediction of any customer’s bill.

ProviderModelInputOutputImportant condition
OpenAIGPT-6 Luna$0.10$0.50Long context: $0.20 / $0.75. Standard cached input: $0.01.[1]
OpenAIGPT-6 Sol$2.00$10.00Long context: $4.00 / $15.00. Standard cached input: $0.20.[1]
OpenAIGPT-6 Astra$10.00$50.00Premium tier, as shown in the supplied chart.[1]
AnthropicClaude Sonnet 5$2.00$10.00Published API rate.[2]
AnthropicClaude Opus 5.5$4.00$20.00Published API rate.[3]
AnthropicClaude Fable 5.1$10.00$50.00Premium tier, as shown in the supplied chart.[2]
xAIGrok 4.7$2.00$6.00Below 200K input tokens. At or above 200K: $4.00 / $12.00.[4]
DeepSeekV4.1 Flash$0.15 off peak, $0.30 peak$0.60 off peak, $1.20 peakCache-hit input: $0.003 off peak, $0.006 peak.[5]
Z.AIGLM-5.3-Flash$0.15$0.50Lower-cost Flash tier.[8]
Z.AIGLM-5.3$1.40$4.40Larger model tier.[8]
MoonshotKimi K2.6$0.95 uncached$4.00Cached input: $0.16.[7]
MoonshotKimi K3$3.00$15.00One-million-token context window.[7]
Alibaba CloudQwen3-MaxCNY 2.50CNY 10.00Up to 32K input tokens; higher context tiers cost more.[6]
Alibaba CloudQwen3.8-MaxCNY 12.00CNY 36.00Model Studio pricing; deployment region matters.[6]
MiniMaxMiniMax-M3CNY 2.10CNY 8.40Up to 512K input tokens; listed as a permanent 50% reduction.[9]

How to read the table

The $2 tier is real, but narrow. It describes selected US models at particular context lengths and cache states. It is not a global market floor. DeepSeek and GLM list dollar-denominated input rates below it, while Kimi spans from a sub-dollar input tier to a more expensive flagship.[5][7][8]

Context and caching change the price. OpenAI’s long-context rates are higher than its standard rates. DeepSeek’s price changes by time of day and whether input tokens hit cache. Grok’s rate doubles at the 200K threshold.[1][4][5]

Keep currencies and regions visible. Alibaba and MiniMax publish the cited rates in yuan. Converting them to dollars would not make endpoints, account eligibility, taxes or service conditions equivalent. Compare like with like before treating one row as cheaper.[6][9]

The useful number is cost per result

A token tariff cannot tell you which model is cheapest for your workload. A lower-priced model may need more retries or a stronger model to repair its output. Measure total input, output, cache use and retries, then divide by tasks that pass the same verifier.

The screenshot’s numbers are a useful clipping of the market. This table shows why a model name and two prices are not enough: the endpoint, context, cache, time window and region are part of the price.

Sources

[1] https://developers.openai.com/api/docs/pricing [2] https://www.anthropic.com/news/claude-sonnet-5 [3] https://www.anthropic.com/claude-opus-5-5 [4] https://docs.x.ai/developers/pricing.md [5] https://api-docs.deepseek.com/quick_start/pricing [6] https://help.aliyun.com/en/model-studio/model-pricing [7] https://platform.kimi.ai/docs/pricing/chat [8] https://docs.z.ai/guides/overview/pricing [9] https://platform.minimaxi.com/docs/guides/pricing-paygo

Keep reading