Model Watch

Frontier model benchmarks, pricing, and speed — charted. Short, current, no hype.
Updated 2026-08-17 · window Jun–Aug 2026
Every figure re-pulled from source on the date above.

Max / high reasoning where applicable. Numbers from Artificial Analysis, official benchmark boards, and vendor cards. Every charted number was re-pulled on the date shown.

61 Grok 4.6 · AA Index ▲ 5 vs prev gen
60 Kimi K3 · AA Index ▲ 15 vs prev gen
63 Opus 5 · AA Index ▲ 6 vs prev gen
61 GPT-5.6 Sol · AA Index ▲ 5 vs prev gen
48 Gemini 3.1 Pro · AA Index ▲ 7 vs prev gen
57 Muse Spark 1.2 · AA Index ▲ 4 vs prev gen
How to read these charts
  • Index version All intelligence scores are Artificial Analysis Intelligence Index v4.1.1, re-pulled 2026-08-15. AA's v4.1 to v4.1.1 rescore lifted most models 2-4 points, so these numbers read a notch higher than last month's and are comparable only to each other.
  • Pairs The intelligence chart pairs each lab's newest model with the one it replaced — pale bar = replaced, bright bar = current — so you see both the level and the rate of gain per month.
  • One board Each capability benchmark charts one official board only (tbench.ai / deepswe.datacurve.ai) or the vendor's own reported score — never a mix of harnesses, and both bars in a pair share one harness.
  • Google Google is charted on its Pro line only. Its Flash models are a different tier and are never swapped in to stand for a Pro — but Flash now outscores Pro, so 3.7 Flash rides the price/speed charts and the story is in the flags.
  • Local models The last section is new: openly-downloadable models a home user can run on their own PC, charted by the video-memory needed to run them and by open-weight intelligence.
Verified this update
  • 2026-08-17 Weekly freshness check: no new flagship shipped from the six tracked labs since 2026-08-15, Artificial Analysis is still on Intelligence Index v4.1.1, and Opus 5 still leads at 63 — so the roster held and only the date was advanced.
  • 2026-08-15 Full six-lab re-pull from Artificial Analysis in one sitting on the new v4.1.1 index. Independent second reads confirmed every intelligence score, so the whole roster is on one comparable scale.
  • Grok 4.6 xAI shipped Grok 4.6 on 2026-08-12 at AA index 61 — level with GPT-5.6 Sol and just behind Opus 5 (63) — for the same $2 / $6 pricing as 4.5, the cheapest frontier model on the page.
  • Muse Spark 1.2 Meta's newest is now Muse Spark 1.2 (57), released 2026-08-05; it replaced 1.1 in the two-model window and 1.0 dropped off.
  • DeepSWE board The official DeepSWE board updated 2026-08-13 and now scores Opus 5 (74), GPT-5.6 Sol (73) and Grok 4.6 (67) for the first time — so several labs finally show a real coding-agent bar.
  • Gemini 3.7 Flash Google's newest Flash (3.7) is added to the price, speed and context charts: $0.75 / $3.75 per million and 340 tokens/sec, by far the cheapest and fastest model here.
Open flags & carried-forward numbers
  • Google's flagship Gemini 3.5 Pro has now missed three announced dates and is still unreleased, so Google's newest Pro here remains February's 3.1 Pro (48) — and both current Flash models (3.7 = 56, 3.6 = 52) now outscore it on the index.
  • Too new for boards Grok 4.6 and Gemini 3.7 Flash are not on the Terminal-Bench 2.1 board yet (it last updated 2026-08-06, before they shipped), and neither xAI nor Google published a GPQA Diamond score for them — so those bars are absent rather than guessed.
  • Muse Spark 1.2 Muse Spark 1.2 posts a DeepSWE coding score (55) but is not on the Terminal-Bench board or reporting a GPQA number yet, and AA has not published its speed or latency — so it is missing from those charts only.
  • Cost-per-task pulled The old cost-per-task chart was removed this update: its numbers predated the v4.1.1 and Grok 4.6 refresh and would have mislabeled Grok 4.5's cost as 4.6's. It returns once Artificial Analysis's published cost-to-run-the-index basis is charted.
  • Auto-updater This page had frozen at 2026-07-27 because its weekly updater was retired in the 2026-08-04 scheduler purge. The refresh capability has been rebuilt; re-arming it on a schedule is a live decision for Phil.
  • Local VRAM Video-memory figures in the local section are 4-bit estimates derived from verified parameter counts, not benchmarked measurements, and Gemma 4 and Qwen3.8-27B are not on AA's index yet so no score is shown for them.

Overall intelligence — each lab's newest model vs the one it replaced

Artificial Analysis Intelligence Index · v4.1.1 · single snapshot 2026-08-15

  • Bars Pale = the model it replaced, bright = current. The block over each pair shows the gain and how fast it was earned (points per month).
  • Anthropic Opus 5 (63) now edges its pricier sibling Fable 5 (62) on the index at half the price; the pair charts Opus 4.8 to Opus 5.
  • Google Google's bar is the oldest here because Gemini 3.5 Pro is still unreleased after three missed dates — February's 3.1 Pro is its newest Pro, and its own Flash line now outscores it.
  • Meta Meta's newest is Muse Spark 1.2 (57), released 2026-08-05, which replaced 1.1.

Capability benchmarks (higher is better)

Score % · pale = previous model, bright = current · same test, same official board, same harness — or no bar at all

  • Terminal-Bench 2.1 Official tbench.ai board, last updated 2026-08-06 — before Grok 4.6, Muse Spark 1.2 and Gemini 3.7 Flash shipped, so those newest models are absent and their previous-gen bars stand alone.
  • DeepSWE Official deepswe.datacurve.ai board v1.1, updated 2026-08-13: it now scores Opus 5 (74), GPT-5.6 Sol (73), Grok 4.6 (67) and Muse Spark 1.2 (55) for the first time, so most labs finally show a real coding-agent pair. Kimi K2.6 and Gemini 3.0 Pro are the only roster models still unscored here.
  • GPQA Diamond Vendor-reported only. xAI published no GPQA for Grok 4.5 or 4.6, so xAI has no GPQA bar; Muse Spark 1.2's is not out yet, so 1.1 shows alone.
Price · USD · lower is better

Token rates

USD per 1M tokens · Gemini 3.7 Flash is the cheapest here; cache hits $0.075

Speed & latency

Output speed (higher is better)

Tokens per second · Gemini 3.7 Flash added · Muse Spark 1.2 not yet measured by AA

Time to first token (lower is better)

Seconds · includes thinking time for reasoning models

Context

Context window (higher is better)

Thousands of tokens

Local models for home users

Local models for home users — will it run on your GPU?

Estimated video memory to run at 4-bit, in GB · green fits a 12-16GB card, blue needs ~24GB

  • What this is Models you download and run on your own machine — private, no API bill, works offline. The real limit is video memory (VRAM), so this chart is what each one needs at 4-bit.
  • 12-16GB card Best pick is Google's Gemma 4 12B (~8GB, Apache-2.0, ships ready-to-run GGUF); on a 16GB card gpt-oss-20b is the stronger reasoner. These are the green bars.
  • 24GB card Qwen3.8-27B is the pick — the best all-round dense model at this size (~16GB, Apache-2.0). For more speed, the MoE options (Nemotron 3.5 Lightning, GLM-4.7-Flash) or Meta's agent-focused Muse Glimmer 30B.
  • Big Mac / lots of RAM With 96-128GB of unified memory, gpt-oss-120b is the standout (~63GB, Apache-2.0); a 256GB Mac Studio can run DeepSeek V4 Flash (MIT) for near-frontier open intelligence.
  • Honesty These GB figures are 4-bit estimates from verified parameter counts, not measured. An AMD RX 6700 XT is a dead-end for local LLMs — assume NVIDIA, an Apple chip, or CPU+RAM.

The open-weight ceiling — AA Intelligence Index

Same v4.1.1 index as the charts above · the strongest models you can legally download

  • Data-center only Kimi K3 (60) is the highest-scoring open-weight model in existence and DeepSeek V4 Pro / GLM-5.2 (both 53) are close behind — all free to download under permissive licenses, but each needs data-center hardware, not a home PC.
  • What you can run at home On a single consumer GPU the best open scores are lower — Muse Glimmer 30B around 35 — and the two best small dense models, Gemma 4 and Qwen3.8-27B, are not on AA's index yet, so run them for quality-per-GB rather than a headline number.

Who leads each charted metric

  • Fastest real iteration (same model line) Moonshot leads the window — Kimi K2.6 to K3 was +15 points (44.9 to 60) in under three months. xAI's Grok 4.5 to 4.6 (+5 in about five weeks) and Meta's Muse Spark 1.1 to 1.2 (+4 in about five weeks) are the fastest cadences; Anthropic's Opus 4.8 to Opus 5 added +6. The chart derives every rate from the release dates, so the points-per-month figures cannot drift from the bars.
  • Grok 4.6 returns xAI to the frontier — at half the price Grok 4.6 (2026-08-12) hit AA index 61, level with GPT-5.6 Sol and just behind Opus 5 (63) and Fable 5 (62), while charging $2 in / $6 out — roughly half what the other frontier models cost. It kept Grok 4.5's 500K context and posts a 67 on the newly-updated DeepSWE coding board. xAI did not publish a GPQA Diamond score, so that bar is absent rather than borrowed from a different harness.
  • Google's Flash now beats Google's Pro The oddest result on the page: Gemini 3.7 Flash scores 56 and 3.6 Flash 52 on the intelligence index, both above Google's newest Pro, the February 3.1 Pro at 48. Gemini 3.5 Pro has slipped past three dates and is still unreleased. Flash is also the cheapest ($0.75 / $3.75) and fastest (340 tokens/sec) model here. The intelligence chart still holds Google to its Pro line so the cross-lab comparison stays Pro-vs-Pro, but on price and speed Flash earns its own bars.
  • Current-gen leaders (AA Index v4.1.1) Opus 5 63 · Fable 5 62 · Grok 4.6 61 · GPT-5.6 Sol 61 · Kimi K3 60 · Muse Spark 1.2 57 · Gemini 3.1 Pro 48. Grok 4.6 and GPT-5.6 Sol are tied, and Kimi K3 sits one point back as the strongest openly-downloadable model.
  • Kimi K3 is the open-weight ceiling Moonshot's Kimi K3 (2.8 trillion parameters, index 60) is the highest-scoring model anyone can download — essentially frontier-level intelligence with public weights. The catch is hardware: it needs a data center to run, so it sets the ceiling for what open weights can do, not what a home user can host. The best you can actually run on one consumer card scores far lower.
  • A benchmark score is not a verdict on whether a model can do YOUR job Worth saying plainly, because these charts invite the opposite read. Gemini 3.1 Pro scores 11.8 on DeepSWE — dead last here by a wide margin — yet it does real day-to-day coding perfectly well inside an agentic IDE. A benchmark measures one harness, one prompt style, one task shape; change any of them and the ranking moves. Read these bars as evidence about a specific test, never as a ceiling on what a model can do for you.
Sources: Artificial Analysis per-model pages — all six labs re-pulled together 2026-08-15 on Intelligence Index v4.1.1 (Grok 4.6 high, Kimi K3 max, Opus 5 max, GPT-5.6 Sol max, Gemini 3.1 Pro Preview, Muse Spark 1.2 xhigh; plus Gemini 3.7 Flash) for index, price, speed, latency and context, with a second independent read confirming every intelligence score · tbench.ai Terminal-Bench 2.1 official board (last updated 2026-08-06) · deepswe.datacurve.ai board v1.1 (updated 2026-08-13) · GPQA Diamond from vendor launch materials · Hugging Face model cards and vendor pricing pages for the local-model section. UNVERIFIED and therefore not charted: GPQA Diamond for Grok 4.6 and Gemini 3.7 Flash (not vendor-reported); Muse Spark 1.2 speed/latency (not yet on AA).
Data lives in src/data/model-watch.json — charts regenerate from it automatically.