Preview build — figures marked “✓” or in green are sourced (benchmarks by Artificial Analysis); treat everything else (value scores, “Est.” figures, unmarked prices and changes) as illustrative. Methodology
Methodology
Where the numbers come from
This page is generated from the same dataset the site renders (snapshot 2026-08-01), so it cannot drift from what you see in the tables. It states plainly what is sourced, what is still an illustrative sample, and how the two are kept apart.
What is sourced today
- 15 of 15 tracked LLM list prices are sourced from public price feeds — 15 confirmed by at least two independent feeds, 0 from a single feed (marked accordingly). Sourced rows show a ✓ src tag with the source list and retrieval date on hover.
- 14 of those are additionally confirmed against the provider's own pricing page (noted in the row's hover text with the check date).
- 32 benchmark scores are sourced measurements from Artificial Analysis (Intelligence Index for language models; arena Elo with 95% confidence intervals for video, TTS and image). A model absent from their data shows “—” and is excluded from the chart, frontier and value ranking — a gap is more honest than mixing measured and invented values on one axis.
- 23 media list prices are sourced: image and OpenAI TTS from LiteLLM (single-feed for now), video and the remaining TTS voices from each provider's own pages with evidence quotes in
data/verification/. Media pricing is tiered by resolution, audio and voice class — the tier each figure applies to is shown on hover. - Still illustrative: all value scores, every price-change entry prefixed “Illustrative sample”, and the price-history chart. Any future row whose price cannot be sourced will carry an est tag — none currently do.
- Unit conversions — some providers publish a price in a different unit than the category uses (e.g. image models billed per token). Those prices are converted, never estimated, and the arithmetic is published on the row and its model page: nano-banana-2 (1,120 output image tokens × $60 per 1M = $0.067 per 1024×1024 image (Google publishes both figures)); nano-banana-pro (1,120 output image tokens × $120 per 1M = $0.134 per image, 1K–2K (Google publishes both figures)).
- Some sourced prices are base-tier prices: Gemini 3 Pro Preview (base tier ≤200K context — above 200K: $4/$18 per 1M); Gemini 2.5 Pro (base tier ≤200K context — above 200K: $2.5/$15 per 1M). The tier caveat is shown on those rows on hover.
Sources
sha256 5cbf1504d795e09709c06d100a08ae80b63efef72bc9f415354a7b8e212ef43c
sha256 ba37bb46dc4662f4bfed71db90168644ef917713fa073bd397775376d55a6b1c
sha256 dcc2f8c9c377520b7e06b5ff49b5dc5344ee5c9f6985c115f6077cfb0c3b33e6
sha256 b8270e942a9800d62386f485bf370bfd70187f29c09010c4862705003015e191
sha256 7eb5c2781642acc3d6c0b590b2648d6ef4f020dd41be08b5a29af62db262d993
sha256 a238547a62284729ea4fda62ad8cac2a6b6a832297b37c040f1e251e67961e72
sha256 2457ffc6d42199586a0a3291ef7dc104b92907aa2a8e4979813071ca2b17a0d0
sha256 8d95326e039cedb1ddf3425601a793677102477c311f0dc98922d8efe13beec1
Snapshots are pulled by scripts/fetch-prices.mjs and normalised by scripts/build-catalog.mjs; both are in the repository, and the mapping from every model row to its exact feed entry is an explicit, reviewable table in that script. On top of the feeds, prices are checked against each provider's own pricing page, with the evidence (page URL, check date, quoted price lines) stored under data/verification/.
Cross-check policy
- models.dev is the primary display source; LiteLLM and Helicone corroborate. A price leg where two or more feeds agree within 0.5% is agree; one feed only is single-source.
- The provider's own pricing page is the final authority. A page confirmation is recorded on the row; if the page confirms a value one feed disputed, the dissenting feed is outvoted — on record, not silently. If the page contradicts the feeds, the offering is demoted to sample until resolved.
- If feeds disagree on a displayed price leg (input or output), the offering is conflict: no sourced price is shown, the row stays a sample, and all conflicting values are recorded for review.
- Feeds also disagree on 2 price legs the site does not yet display: cachedInput of gemini-2-5-flash (models.dev $0.03 vs LiteLLM $0.03 vs Helicone $0.075); cachedInput of gemini-2-5-pro (models.dev $0.125 vs LiteLLM $0.125 vs Helicone $0.31). These do not affect any shown price, but they are recorded rather than hidden.
- A model with no unambiguous identity in any feed is not guessed at. 0 offerings are currently in this state (e.g. sample models that no live feed carries) and remain samples.
- Offerings the provider no longer publishes a price for are retired from the catalogue, on record: DeepSeek R1 (DeepSeek, Aug 1, 2026 — no longer on DeepSeek's official price list (superseded by the V4 lineup)); DeepSeek V3.2 (DeepSeek, Aug 1, 2026 — no longer on DeepSeek's official price list (superseded by the V4 lineup)); Mistral Medium 3 (Mistral, Aug 1, 2026 — superseded by Mistral Medium 3.5 on Mistral's price page; the 3.0 rate is no longer published); Llama 4 Maverick (Fireworks, Aug 1, 2026 — moved off Fireworks serverless per-token pricing (on-demand GPU deployment only)); Llama 4 70B (Together / Fireworks / Groq, Aug 2, 2026 — no such model exists in any connected feed or provider catalogue (the Llama 4 line ships as Scout/Maverick 17B variants) — unsourceable); DeepSeek R1 (Groq, Aug 2, 2026 — Groq's catalogue does not list DeepSeek R1 in any connected feed — unsourceable); Qwen 3 235B (Fireworks, Aug 2, 2026 — Fireworks lists no Qwen3-235B serverless offering in any feed or on its pricing pages — unsourceable); GPT-5.2 mini (OpenAI, Aug 2, 2026 — absent from all connected feeds and OpenAI's pricing page under any plausible identifier — unsourceable); Mistral Small 3 (Mistral, Aug 2, 2026 — Mistral's price list carries Small 3.2 and Small 4 only; the 3.0 rate is no longer published — unsourceable).
The value score — and its limits
- The estimated value score is benchmark-per-dollar relative to the category median, log-normalised to 0–100 across the whole catalogue. Its benchmark inputs are now sourced (Artificial Analysis), but video / TTS / image price inputs are still samples and the normalisation below is provisional — which is why every value score keeps the “Est.” label.
- LLM price sorting blends input and output at 3:1. That is an editorial assumption — real workloads range from ~20:1 (retrieval-heavy) to ~1:3 (generation-heavy) — and it will become a user-adjustable setting.
- Known limitation: the normalisation is computed over the live catalogue, so adding a model shifts other scores. A versioned, frozen scoring methodology (plus a Pareto-frontier view that needs no score at all) is planned to replace it.
Use the data
The full dataset behind every table is downloadable — the same serializer feeds the site, the API and the CSV, so they cannot disagree:
/api/v1/models— JSON with per-row provenance (sources, retrieval dates, provider-verification links, tier notes)./api/v1/models.csv— the same rows as CSV, provenance columns included.
License: the compilation is CC BY 4.0 — use it, cite “Model Price Book, snapshot 2026-08-01”. Benchmark scores are Artificial Analysis data used with required attribution: if you redistribute them, keep the Artificial Analysis credit. Synthetic values are never exported — the est. value score ships only under its value_score_est name with an explicit caveat in the payload notes.
Corrections
Spotted a wrong number? Email corrections@modelpricebook.com. Every correction we make is published with before and after values, disputed datapoints are marked under review rather than silently edited, and unsourceable models are retired on record — see the corrections log.
Catalog generated Aug 2, 2026 · snapshot 2026-08-01 · all prices USD, provider list price, excluding taxes, caching, batch discounts, volume tiers and (except where noted per row) context-length tiers.