Claude vs GPT vs Grok vs DeepSeek: A Price-Based Comparison (August 2026)
Benchmark leaderboards for large language models contradict each other badly enough to be unusable. Published prices do not. This is a structured comparison built on the numbers each lab commits to contractually.
Why not compare benchmark scores?
Three reasons, in order of severity.
Leaderboards disagree by more than the gaps they report. Four separate sites gave four different figures for the same model on the same benchmark, spanning roughly eight percentage points. On a 500-problem set that is forty problems.
Some published numbers cannot be real. One site listed a superseded model as current flagship. Another quoted a precise score for a model available only by invitation.
The score belongs to the scaffold, not the model. On agentic coding benchmarks the harness determines much of the result. The SWE-bench team published an agent in 100 lines of Python that scores 65% on SWE-bench Verified by itself. Two models tested under two different harnesses are not comparable, which describes almost every roundup.
Takeaway: treat any cross-vendor benchmark table without a stated, identical harness as unusable.
What do the models cost?
| Model | Lab | Input | Output | Cached input | Context |
|---|---|---|---|---|---|
| Claude Fable 5 | Anthropic | $10.00 | $50.00 | $1.00 | 1M |
| GPT-5.6 Sol | OpenAI | $5.00 | $30.00 | $0.50 | ~270K+ |
| Claude Opus 5 | Anthropic | $5.00 | $25.00 | $0.50 | 1M |
| GPT-5.6 Terra | OpenAI | $2.50 | $15.00 | $0.25 | ~270K+ |
| Claude Sonnet 5 | Anthropic | $2.00 | $10.00 | $0.20 | 1M |
| Grok 4.5 | xAI | $2.00 | $6.00 | not published | 500K |
| GPT-5.6 Luna | OpenAI | $1.00 | $6.00 | $0.10 | ~270K+ |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 | $0.10 | 200K |
| DeepSeek V4-Pro | DeepSeek | $0.435 | $0.87 | $0.003625 | 1M |
| DeepSeek V4-Flash | DeepSeek | $0.14 | $0.28 | $0.0028 | 1M |
All figures are USD per million tokens. Sonnet 5 shows introductory pricing through 31 August 2026; standard is $3/$15 from 1 September.
Output-price spread, top to bottom: 179x.
Takeaway: no capability claim anywhere approaches a 179x gap.
What does one coding session cost?
Modelled on forty agent turns, 15,000 input tokens and 1,200 output tokens per turn, 90% cache hit rate.
| Model | No cache | With cache | vs cheapest |
|---|---|---|---|
| Claude Fable 5 | $8.400 | $3.540 | 152x |
| GPT-5.6 Sol | $4.440 | $2.010 | 86x |
| Claude Opus 5 | $4.200 | $1.770 | 76x |
| GPT-5.6 Terra | $2.220 | $1.005 | 43x |
| Claude Sonnet 5 | $1.680 | $0.708 | 30x |
| Grok 4.5 | $1.488 | $1.488 | 64x |
| GPT-5.6 Luna | $0.888 | $0.402 | 17x |
| Claude Haiku 4.5 | $0.840 | $0.354 | 15x |
| DeepSeek V4-Pro | $0.303 | $0.070 | 3x |
| DeepSeek V4-Flash | $0.097 | $0.023 | 1x |
Two corrections that matter:
Grok's cached column is overstated. xAI publishes no cache-read rate in its model table, so the figure repeats the uncached cost.
Claude's real cost is higher than shown. Anthropic's documentation states that models from 4.7 onward use a newer tokenizer producing roughly 30% more tokens for the same text. Adjusted, the Opus 5 session lands nearer $5.46 uncached rather than $4.20.
Takeaway: compare cache terms and tokenizers, not sticker prices.
Where do the Chinese labs actually lead?
Not on the leaderboard. On the economics, and on openness.
Price floor. DeepSeek V4-Pro delivers a 1M context and 384K max output at roughly one twenty-ninth of Opus 5's output price. Cache hits cost $0.003625 per million, effectively free.
Distribution through a competitor's tooling. DeepSeek publishes an Anthropic-format endpoint at api.deepseek.com/anthropic and documents wiring their models into Claude Code. That is a deliberate strategy: if the client is standardised, the model is swappable.
Capacity pressure as a signal. DeepSeek is introducing peak/off-peak pricing at 2x during Beijing business hours. Labs do not ration what nobody wants.
Scale and licensing. Moonshot's Kimi K3, released 17 July 2026, is billed as the largest open-source model at 2.8 trillion parameters. Alibaba's Qwen family ships the widest range under permissive licences.
The specific benchmark claims made for these models could not be verified here, for the reasons in the first section. What is verifiable is pricing, context, licensing, and API surface. On those, the gap is not close.
Takeaway: the Chinese labs removed price as a reason to compromise. That has more consequences than a leaderboard position.
Which model should you use?
There is no general answer. There are four situational ones.
Long agentic runs where failure is expensive. Premium tier is defensible. A 14x price gap needs the expensive model to be better often enough that avoided retries and review time exceed roughly four dollars a session. For a payments refactor that is easy arithmetic.
High-volume, low-stakes work. Classification, extraction, summarisation, first drafts. The cheap tier is the correct answer, not a compromise. At three cents a session the argument is over.
Data that cannot leave your infrastructure. Open weights are the only option at any price, which makes capability ranking secondary.
Everything else. Route rather than choose. Cheap model first, escalate on failure. Cheap to build, and it converts the entire question into an implementation detail.
Takeaway: the decision is about your failure cost, not about a ranking.
The honest read
Prices are not capability. The 179x gap describes what you pay, not what you get. A cheap model that fails on the 5% of tasks that matter most is not cheap.
The counter-argument is real. Frontier labs solve the hard problems first, and their safety and reliability work is more mature. That is a genuine product, even though it is harder to defend than a benchmark lead because it erodes quietly rather than being beaten publicly.
Session costs are a model, not a quote. Turn counts and token volumes are estimates; substitute your own.
Prices move fast. Every figure here has a date on it for that reason. Anthropic's Sonnet 5 pricing changes on 1 September 2026 alone.
One correction made during research. An unfamiliar model name in a leaderboard was initially assumed to be fabricated. Checking the vendor's documentation showed it was real, in a limited-availability programme. Scepticism about the category was warranted; the specific judgement was wrong.

Written by
The Builder’s PlaybookEssays on building for the web in the AI era — engineering careers, developer economics, and the workflows behind shipped products. Written by the team at ReactBD.
View profileKeep reading
More from ReactBD
Claude Skills: What Each One Costs, and How Many You Can Run
Every Claude skill you install occupies context on every message you send, whether or not it fires. This is a short overview of what that costs, which skills are worth the space, and where the ceiling sits.