Research reportai-agents
Terminal-Bench 2.1: How Our Indexed Coding Agents Actually Score
Stanford/Laude Institute's independent Terminal-Bench results, matched against the coding agents in our index — accuracy, model pairing and cost, not vendor claims.
Every data report we’ve published so far about our coding-agents category has been about the business layer: licensing and portability, pricing, which model each tool won’t tell you it runs, GitHub star velocity. None of it answers the question a developer picking a tool actually asks first: which one is better at the job? We don’t run our own coding benchmark, and a vendor’s own cherry-picked number isn’t evidence — so instead we went to the one dataset that is both independent and reproducible: Terminal-Bench, built by Stanford University and the Laude Institute (with Snorkel AI co-authoring the 2.0 revision), which scores agent-plus-model pairings on real terminal coding tasks under identical, disclosed conditions. We pulled its public leaderboard directly on 2026-08-02 and matched every agent row against our own index.
The current leaderboard: Terminal-Bench 2.1
2.1 is Terminal-Bench’s live, current leaderboard — 17 entries as of this check, the newest model generations only. Four of our indexed coding agents appear on it directly:
| Agent (our listing) | Model | Accuracy | Cost per run |
|---|---|---|---|
| Claude Code | Claude Fable 5 | 83.8% ± 1.2% | $552.67 |
| OpenAI Codex CLI | GPT-5.5 | 83.1% ± 1.1% | $2,059.19 |
| Cursor CLI (Cursor’s terminal agent) | Grok 4.5 | 79.3% ± 1.5% | $134.09 |
| Claude Code | Claude Opus 4.8 | 78.9% ± 1.3% | $286.94 |
| OpenAI Codex CLI | GPT-5.6 Terra | 78.4% ± 1.3% | $421.15 |
| OpenAI Codex CLI | GPT-5.6 Luna | 75.7% ± 1.3% | $241.45 |
| Claude Code | Claude Sonnet 5 | 74.6% ± 1.6% | $288.18 |
| Claude Code | Claude Opus 4.7 | 68.9% ± 1.4% | $599.52 |
| Gemini CLI (Antigravity CLI’s predecessor) | Gemini 3 Pro | 65.8% ± 1.4% | $247.76 |
| Gemini CLI (Antigravity CLI’s predecessor) | Gemini 3.1 Pro | 65.8% ± 1.7% | $236.49 |
| Claude Code | GLM-5.1 | 58.7% ± 1.2% | $277.14 |
Two things worth reading past the headline rank. First, accuracy is a property of the pairing, not the harness alone — Claude Code spans a 25-point range (58.7% to 83.8%) purely on which model it’s pointed at, so “Claude Code scores 83.8%” is only true for the specific Fable 5 configuration Anthropic’s own harness was tested with; the same tool paired with GLM-5.1 scores worse than every other agent on this table. Second, cost and accuracy don’t move together: the top-scoring Codex CLI run cost $2,059.19 for the full evaluation, roughly 4x Claude Code’s $552.67 top run for a 0.7-point accuracy difference, while Cursor’s terminal agent posted the third-highest accuracy (79.3%) at under a quarter of Claude Code’s top-run cost ($134.09).
The broader field: Terminal-Bench 2.0
2.0 is the previous, larger leaderboard (142 entries spanning an earlier round of models) and it’s where four more of our indexed tools have independent results that 2.1 hasn’t re-tested yet — read these as a genuine but older data point, not directly comparable score-for-score against the 2.1 table above, since both the task set’s difficulty calibration and the model generations differ between versions:
| Agent (our listing) | Best model pairing | Accuracy |
|---|---|---|
| Factory (“Droid”) | GPT-5.3-Codex | 77.3% ± 2.2% |
| Warp | (multiple models tested) | 61.2% ± 3.0% |
| OpenHands | Claude Opus 4.5 | 51.9% ± 2.9% |
| OpenCode | Claude Opus 4.5 | 51.7% ± N/A |
Factory’s agent product, Droid, is the standout here — a 77.3% best score on 2.0 would have placed 6th on the current 2.1 table, ahead of every Claude Code pairing except the top two. It hasn’t yet appeared on 2.1, so we can’t confirm whether that holds against the newest model generation, but the 2.0 result alone is real, independent evidence, not a vendor claim.
What’s absent, and why that’s worth saying plainly
Aider is one of the six tools in our own terminal-agent portability piece and is exactly the kind of bring-your-own-model terminal harness Terminal-Bench tests — but it has no entry on either leaderboard version. Aider publishes its own benchmark instead of submitting to independent evaluations like this one, which is a materially different evidentiary standard: a tool grading its own homework isn’t the same claim as a third-party academic lab running it under disclosed, identical conditions. We’re not asserting Aider performs worse — we simply have no independent number for it, and say so rather than leaving the gap silent.
Antigravity CLI — the tool that replaced Gemini CLI’s free/personal tier on 2026-06-18, as we’ve covered before — has no Terminal-Bench entry of its own yet either. The “Gemini CLI” rows in both tables above are the predecessor product under its old name, included for context only; they are not a score for Antigravity CLI, which is a distinct, closed-source successor and hasn’t been independently benchmarked under its own name as of this check.
The rest of our 19-listing coding-agents category — GitHub Copilot, Windsurf, Cline, Zed, Devin, Replit Agent, Bolt.new, CodeRabbit, Juggler and Zot — don’t appear on Terminal-Bench at all. That’s consistent with what the benchmark actually measures (autonomous terminal task completion), not a verdict on those tools: an IDE-embedded assistant, a browser-based app builder, or a PR-review bot isn’t the shape of product this specific benchmark was built to score.
Methodology
Every number above was read directly from Terminal-Bench’s own public leaderboard pages — tbench.ai/leaderboard/terminal-bench/2.1 and …/2.0 — accessed 2026-08-02, not from a third-party summary; we cross-checked the 2.1 page twice on independent fetches and got identical figures both times. Terminal-Bench is built and run by Stanford University and the Laude Institute (Snorkel AI co-authored the 2.0 revision) — an independent academic benchmark, not a vendor-operated one. We matched leaderboard “agent” names to our own listings by product identity, not string similarity: Factory’s row appears as “Droid” (its product’s actual name, confirmed against our own listing copy) rather than “Factory”; Cursor’s row is its separate CLI agent product, not the IDE our main Cursor listing describes. The 2.1 table reflects the “Displaying 17 of 17 available entries” full leaderboard as shown at access time, with individual row submission dates ranging 2026-05-01 to 2026-07-11. Confidence intervals, effort tier and cost-per-run figures are transcribed as published; we did not recompute them. We are not affiliated with Terminal-Bench, Stanford, or the Laude Institute, and none of the agents above paid for placement in this piece — that’s the entire reason we used a third-party benchmark instead of a self-reported one.