AI Tools

Claude vs GPT vs Grok vs DeepSeek: A Price-Based Comparison (August 2026)

Benchmark leaderboards for large language models contradict each other badly enough to be unusable. Published prices do not. This is a structured comparison built on the numbers each lab commits to contractually.

Why not compare benchmark scores?

Three reasons, in order of severity.

Leaderboards disagree by more than the gaps they report. Four separate sites gave four different figures for the same model on the same benchmark, spanning roughly eight percentage points. On a 500-problem set that is forty problems.

Some published numbers cannot be real. One site listed a superseded model as current flagship. Another quoted a precise score for a model available only by invitation.

The score belongs to the scaffold, not the model. On agentic coding benchmarks the harness determines much of the result. The SWE-bench team published an agent in 100 lines of Python that scores 65% on SWE-bench Verified by itself. Two models tested under two different harnesses are not comparable, which describes almost every roundup.

Takeaway: treat any cross-vendor benchmark table without a stated, identical harness as unusable.


What do the models cost?

Model Lab Input Output Cached input Context
Claude Fable 5Anthropic$10.00$50.00$1.001M
GPT-5.6 SolOpenAI$5.00$30.00$0.50~270K+
Claude Opus 5Anthropic$5.00$25.00$0.501M
GPT-5.6 TerraOpenAI$2.50$15.00$0.25~270K+
Claude Sonnet 5Anthropic$2.00$10.00$0.201M
Grok 4.5xAI$2.00$6.00not published500K
GPT-5.6 LunaOpenAI$1.00$6.00$0.10~270K+
Claude Haiku 4.5Anthropic$1.00$5.00$0.10200K
DeepSeek V4-ProDeepSeek$0.435$0.87$0.0036251M
DeepSeek V4-FlashDeepSeek$0.14$0.28$0.00281M

All figures are USD per million tokens. Sonnet 5 shows introductory pricing through 31 August 2026; standard is $3/$15 from 1 September.

Output-price spread, top to bottom: 179x.

Takeaway: no capability claim anywhere approaches a 179x gap.


What does one coding session cost?

Modelled on forty agent turns, 15,000 input tokens and 1,200 output tokens per turn, 90% cache hit rate.

Model No cache With cache vs cheapest
Claude Fable 5$8.400$3.540152x
GPT-5.6 Sol$4.440$2.01086x
Claude Opus 5$4.200$1.77076x
GPT-5.6 Terra$2.220$1.00543x
Claude Sonnet 5$1.680$0.70830x
Grok 4.5$1.488$1.48864x
GPT-5.6 Luna$0.888$0.40217x
Claude Haiku 4.5$0.840$0.35415x
DeepSeek V4-Pro$0.303$0.0703x
DeepSeek V4-Flash$0.097$0.0231x

Two corrections that matter:

Grok's cached column is overstated. xAI publishes no cache-read rate in its model table, so the figure repeats the uncached cost.

Claude's real cost is higher than shown. Anthropic's documentation states that models from 4.7 onward use a newer tokenizer producing roughly 30% more tokens for the same text. Adjusted, the Opus 5 session lands nearer $5.46 uncached rather than $4.20.

Takeaway: compare cache terms and tokenizers, not sticker prices.


Where do the Chinese labs actually lead?

Not on the leaderboard. On the economics, and on openness.

Price floor. DeepSeek V4-Pro delivers a 1M context and 384K max output at roughly one twenty-ninth of Opus 5's output price. Cache hits cost $0.003625 per million, effectively free.

Distribution through a competitor's tooling. DeepSeek publishes an Anthropic-format endpoint at api.deepseek.com/anthropic and documents wiring their models into Claude Code. That is a deliberate strategy: if the client is standardised, the model is swappable.

Capacity pressure as a signal. DeepSeek is introducing peak/off-peak pricing at 2x during Beijing business hours. Labs do not ration what nobody wants.

Scale and licensing. Moonshot's Kimi K3, released 17 July 2026, is billed as the largest open-source model at 2.8 trillion parameters. Alibaba's Qwen family ships the widest range under permissive licences.

The specific benchmark claims made for these models could not be verified here, for the reasons in the first section. What is verifiable is pricing, context, licensing, and API surface. On those, the gap is not close.

Takeaway: the Chinese labs removed price as a reason to compromise. That has more consequences than a leaderboard position.


Which model should you use?

There is no general answer. There are four situational ones.

Long agentic runs where failure is expensive. Premium tier is defensible. A 14x price gap needs the expensive model to be better often enough that avoided retries and review time exceed roughly four dollars a session. For a payments refactor that is easy arithmetic.

High-volume, low-stakes work. Classification, extraction, summarisation, first drafts. The cheap tier is the correct answer, not a compromise. At three cents a session the argument is over.

Data that cannot leave your infrastructure. Open weights are the only option at any price, which makes capability ranking secondary.

Everything else. Route rather than choose. Cheap model first, escalate on failure. Cheap to build, and it converts the entire question into an implementation detail.

Takeaway: the decision is about your failure cost, not about a ranking.


The honest read

Prices are not capability. The 179x gap describes what you pay, not what you get. A cheap model that fails on the 5% of tasks that matter most is not cheap.

The counter-argument is real. Frontier labs solve the hard problems first, and their safety and reliability work is more mature. That is a genuine product, even though it is harder to defend than a benchmark lead because it erodes quietly rather than being beaten publicly.

Session costs are a model, not a quote. Turn counts and token volumes are estimates; substitute your own.

Prices move fast. Every figure here has a date on it for that reason. Anthropic's Sonnet 5 pricing changes on 1 September 2026 alone.

One correction made during research. An unfamiliar model name in a leaderboard was initially assumed to be fabricated. Checking the vendor's documentation showed it was real, in a limited-availability programme. Scepticism about the category was warranted; the specific judgement was wrong.

The Builder’s Playbook

Written by

The Builder’s Playbook

Essays on building for the web in the AI era — engineering careers, developer economics, and the workflows behind shipped products. Written by the team at ReactBD.

View profile

Keep reading

More from ReactBD

Claude
6 min

Claude Skills: What Each One Costs, and How Many You Can Run

Every Claude skill you install occupies context on every message you send, whether or not it fires. This is a short overview of what that costs, which skills are worth the space, and where the ceiling sits.