LLM leaderboard
Last updated: 2026-09-25
Picks
Agent-oriented picks.
Index ignores price. Default = best Index within 2× cheapest task $ · Cheap = lowest task $ · Frontier = best Index. Rules under Method.
| Category | 1st | 2nd |
|---|---|---|
| Default | GLM-5.3 | Kimi K3 |
| Cheap | Gemini 3.8 Flash | GPT-6 |
| Frontier | GPT-6 | Claude Opus 5 |
* Fewest active benchmarks on this board.
Leaderboard
Top 15.
Showing the top 15 of 34 models with ≥4 active direct benchmarks. Expand a row for sources and costs.
1 GPT-6
- Hermes Index
- 84.20
- Coverage
- 6/6 direct + Agent Arena
- Value (Index÷task$)
- 19.01
- LMSYS
- 1480 (z=+0.26, w·z=+0.029) · as of 2026-09-25
- Agent Arena
- 97.6% (z=+1.10, w·z=+0.183) · as of 2026-09-25
- Terminal-Bench 4.0
- 58.2% (z=+2.23, w·z=+0.372) · as of 2026-09-25
- BrowseComp
- 91.5% (z=+0.95, w·z=+0.158) · as of 2026-09-25
- HLE
- 54.7% (z=+1.47, w·z=+0.244) · as of 2026-09-25
- GPQA Diamond
- 96.3% ±2.6 (z=+1.69, w·z=+0.094) · as of 2026-09-25
- SWE-Atlas-QnA
- — (absent; imputed z=0)
- DeepSWE
- 74.1% (z=+0.85, w·z=+0.141) · as of 2026-09-25
- DeepSWE $/task
- $4.43
- List $/1M
- $10/50
2 Claude Opus 5
- Hermes Index
- 83.34
- Coverage
- 6/6 direct + Agent Arena
- Value (Index÷task$)
- 7.04
- LMSYS
- 1490 (z=+0.74, w·z=+0.082) · as of 2026-09-25
- Agent Arena
- 95.2% (z=+1.01, w·z=+0.169) · as of 2026-09-25
- Terminal-Bench 4.0
- 53.9% (z=+1.96, w·z=+0.328) · as of 2026-09-25
- BrowseComp
- 90.8% (z=+0.87, w·z=+0.144) · as of 2026-09-25
- HLE
- 54.9% (z=+1.50, w·z=+0.249) · as of 2026-09-25
- GPQA Diamond
- 93.7% ±3.4 (z=+0.48, w·z=+0.027) · as of 2026-09-25
- SWE-Atlas-QnA
- — (absent; imputed z=0)
- DeepSWE
- 73.6% (z=+0.81, w·z=+0.135) · as of 2026-09-25
- DeepSWE $/task
- $11.84
- List $/1M
- $5/25
3 Claude Fable 5.1*
- Hermes Index
- 82.81
- Index band
- 82.81–88.22
- Coverage
- 4/6 direct + Agent Arena
- Value (Index÷task$)
- 1.38
- LMSYS
- 1498 (z=+1.13, w·z=+0.125) · as of 2026-09-25
- Agent Arena
- 100.0% (z=+1.18, w·z=+0.197) · as of 2026-09-25
- Terminal-Bench 4.0
- 57.9% (z=+2.21, w·z=+0.369) · as of 2026-09-25
- BrowseComp
- — (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
- HLE
- 59.1% (z=+2.18, w·z=+0.364) · as of 2026-09-25
- GPQA Diamond
- 93.7% ±3.4 (z=+0.48, w·z=+0.027) · as of 2026-09-25
- SWE-Atlas-QnA
- — (absent; imputed z=0)
- DeepSWE
- — (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
- DeepSWE $/task
- —
- List $/1M
- $10/50
4 Claude Fable 5
- Hermes Index
- 80.83
- Index band
- 80.83–82.60
- Coverage
- 5/6 direct + Agent Arena
- Value (Index÷task$)
- 6.03
- LMSYS
- 1506 (z=+1.46, w·z=+0.162) · as of 2026-09-25
- Agent Arena
- 90.5% (z=+0.84, w·z=+0.141) · as of 2026-09-25
- Terminal-Bench 4.0
- 44.5% (z=+1.38, w·z=+0.229) · as of 2026-09-25
- BrowseComp
- — (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
- HLE
- 55.5% (z=+1.59, w·z=+0.265) · as of 2026-09-25
- GPQA Diamond
- 92.6% ±3.6 (z=-0.05, w·z=-0.003) · as of 2026-09-25
- SWE-Atlas-QnA
- — (absent; imputed z=0)
- DeepSWE
- 69.9% (z=+0.53, w·z=+0.089) · as of 2026-09-25
- DeepSWE $/task
- $13.41
- List $/1M
- $10/50
5 GPT-5.6 Sol
- Hermes Index
- 79.79
- Coverage
- 6/6 direct + Agent Arena
- Value (Index÷task$)
- 9.51
- LMSYS
- 1483 (z=+0.43, w·z=+0.048) · as of 2026-09-25
- Agent Arena
- 85.7% (z=+0.67, w·z=+0.112) · as of 2026-09-25
- Terminal-Bench 4.0
- 37.3% (z=+0.92, w·z=+0.153) · as of 2026-09-25
- BrowseComp
- 92.2% (z=+1.03, w·z=+0.172) · as of 2026-09-25
- HLE
- 49.5% (z=+0.63, w·z=+0.104) · as of 2026-09-25
- GPQA Diamond
- 95.3% ±3.0 (z=+1.20, w·z=+0.067) · as of 2026-09-25
- SWE-Atlas-QnA
- — (absent; imputed z=0)
- DeepSWE
- 72.7% (z=+0.74, w·z=+0.123) · as of 2026-09-25
- DeepSWE $/task
- $8.39
- List $/1M
- $5/30
6 GLM-5.3
- Hermes Index
- 77.24
- Index band
- 77.24–78.29
- Coverage
- 5/6 direct + Agent Arena
- Value (Index÷task$)
- 19.34
- LMSYS
- 1479 (z=+0.23, w·z=+0.026) · as of 2026-09-25
- Agent Arena
- 59.5% (z=-0.25, w·z=-0.042) · as of 2026-09-25
- Terminal-Bench 4.0
- 41.8% (z=+1.21, w·z=+0.201) · as of 2026-09-25
- BrowseComp
- — (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
- HLE
- 55.5% (z=+1.59, w·z=+0.265) · as of 2026-09-25
- GPQA Diamond
- 92.6% ±3.6 (z=-0.05, w·z=-0.003) · as of 2026-09-25
- SWE-Atlas-QnA
- — (absent; imputed z=0)
- DeepSWE
- 69.0% (z=+0.46, w·z=+0.077) · as of 2026-09-25
- DeepSWE $/task
- $3.99
- List $/1M
- $1.4/4.4
7 Kimi K3
- Hermes Index
- 76.18
- Index band
- 76.18–77.01
- Coverage
- 5/6 direct + Agent Arena
- Value (Index÷task$)
- 16.37
- LMSYS
- 1485 (z=+0.49, w·z=+0.054) · as of 2026-09-25
- Agent Arena
- 81.0% (z=+0.51, w·z=+0.084) · as of 2026-09-25
- Terminal-Bench 4.0
- — (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
- BrowseComp
- 91.2% (z=+0.91, w·z=+0.152) · as of 2026-09-25
- HLE
- 46.9% (z=+0.21, w·z=+0.034) · as of 2026-09-25
- GPQA Diamond
- 93.5% ±3.4 (z=+0.39, w·z=+0.021) · as of 2026-09-25
- SWE-Atlas-QnA
- — (absent; imputed z=0)
- DeepSWE
- 68.5% (z=+0.43, w·z=+0.071) · as of 2026-09-25
- DeepSWE $/task
- $4.65
- List $/1M
- $3/15
8 GPT-5.6 Terra
- Hermes Index
- 75.72
- Coverage
- 6/6 direct + Agent Arena
- Value (Index÷task$)
- 15.31
- LMSYS
- 1466 (z=-0.38, w·z=-0.042) · as of 2026-09-25
- Agent Arena
- 47.6% (z=-0.67, w·z=-0.112) · as of 2026-09-25
- Terminal-Bench 4.0
- 21.5% (z=-0.07, w·z=-0.011) · as of 2026-09-25
- BrowseComp
- 87.5% (z=+0.48, w·z=+0.080) · as of 2026-09-25
- HLE
- 58.7% (z=+2.12, w·z=+0.353) · as of 2026-09-25
- GPQA Diamond
- 93.4% ±3.5 (z=+0.34, w·z=+0.019) · as of 2026-09-25
- SWE-Atlas-QnA
- — (absent; imputed z=0)
- DeepSWE
- 69.6% (z=+0.51, w·z=+0.085) · as of 2026-09-25
- DeepSWE $/task
- $4.95
- List $/1M
- $2/12
9 Gemini 3.8 Flash
- Hermes Index
- 75.38
- Index band
- 75.38–76.06
- Coverage
- 5/6 direct + Agent Arena
- Value (Index÷task$)
- 31.91
- LMSYS
- 1493 (z=+0.87, w·z=+0.097) · as of 2026-09-25
- Agent Arena
- 69.0% (z=+0.08, w·z=+0.014) · as of 2026-09-25
- Terminal-Bench 4.0
- 19.1% (z=-0.22, w·z=-0.036) · as of 2026-09-25
- BrowseComp
- — (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
- HLE
- 47.8% (z=+0.36, w·z=+0.059) · as of 2026-09-25
- GPQA Diamond
- 95.3% ±3.0 (z=+1.20, w·z=+0.067) · as of 2026-09-25
- SWE-Atlas-QnA
- — (absent; imputed z=0)
- DeepSWE
- 73.8% (z=+0.83, w·z=+0.138) · as of 2026-09-25
- DeepSWE $/task
- $2.36
- List $/1M
- $0.75/3.75
10 GPT-5.5
- Hermes Index
- 73.93
- Index band
- 73.93–74.32
- Coverage
- 5/6 direct + Agent Arena
- Value (Index÷task$)
- 10.23
- LMSYS
- 1479 (z=+0.22, w·z=+0.024) · as of 2026-09-25
- Agent Arena
- 78.6% (z=+0.42, w·z=+0.070) · as of 2026-09-25
- Terminal-Bench 4.0
- — (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
- BrowseComp
- 84.4% (z=+0.12, w·z=+0.020) · as of 2026-09-25
- HLE
- 45.8% (z=+0.03, w·z=+0.004) · as of 2026-09-25
- GPQA Diamond
- 93.5% ±3.4 (z=+0.39, w·z=+0.021) · as of 2026-09-25
- SWE-Atlas-QnA
- — (absent; imputed z=0)
- DeepSWE
- 67.0% (z=+0.32, w·z=+0.053) · as of 2026-09-25
- DeepSWE $/task
- $7.23
- List $/1M
- $5/30
11 Claude Opus 4.8
- Hermes Index
- 73.87
- Coverage
- 6/6 direct + Agent Arena
- Value (Index÷task$)
- 5.59
- LMSYS
- 1477 (z=+0.14, w·z=+0.016) · as of 2026-09-25
- Agent Arena
- 88.1% (z=+0.76, w·z=+0.127) · as of 2026-09-25
- Terminal-Bench 4.0
- 23.6% (z=+0.07, w·z=+0.011) · as of 2026-09-25
- BrowseComp
- 84.3% (z=+0.11, w·z=+0.018) · as of 2026-09-25
- HLE
- 48.7% (z=+0.49, w·z=+0.082) · as of 2026-09-25
- GPQA Diamond
- 92.0% ±3.8 (z=-0.34, w·z=-0.019) · as of 2026-09-25
- SWE-Atlas-QnA
- — (absent; imputed z=0)
- DeepSWE
- 59.0% (z=-0.29, w·z=-0.048) · as of 2026-09-25
- DeepSWE $/task
- $13.22
- List $/1M
- $5/25
12 Muse Spark*
- Hermes Index
- 73.48
- Index band
- 73.48–74.21
- Coverage
- 4/6 direct + Agent Arena
- Value (Index÷task$)
- 19.88
- LMSYS
- 1493 (z=+0.89, w·z=+0.099) · as of 2026-09-25
- Agent Arena
- 71.4% (z=+0.17, w·z=+0.028) · as of 2026-09-25
- Terminal-Bench 4.0
- — (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
- BrowseComp
- — (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
- HLE
- 48.7% (z=+0.50, w·z=+0.083) · as of 2026-09-25
- GPQA Diamond
- 94.1% ±3.3 (z=+0.67, w·z=+0.037) · as of 2026-09-25
- SWE-Atlas-QnA
- — (absent; imputed z=0)
- DeepSWE
- 54.9% (z=-0.60, w·z=-0.100) · as of 2026-09-25
- DeepSWE $/task
- $3.70
- List $/1M
- $1.25/4.25
13 Claude Sonnet 5
- Hermes Index
- 71.60
- Coverage
- 6/6 direct + Agent Arena
- Value (Index÷task$)
- 2.71
- LMSYS
- 1461 (z=-0.60, w·z=-0.067) · as of 2026-09-25
- Agent Arena
- 83.3% (z=+0.59, w·z=+0.098) · as of 2026-09-25
- Terminal-Bench 4.0
- 12.4% (z=-0.64, w·z=-0.106) · as of 2026-09-25
- BrowseComp
- 84.7% (z=+0.16, w·z=+0.026) · as of 2026-09-25
- HLE
- 48.7% (z=+0.50, w·z=+0.083) · as of 2026-09-25
- GPQA Diamond
- 94.1% ±3.3 (z=+0.67, w·z=+0.037) · as of 2026-09-25
- SWE-Atlas-QnA
- — (absent; imputed z=0)
- DeepSWE
- 53.8% (z=-0.67, w·z=-0.112) · as of 2026-09-25
- DeepSWE $/task
- $26.40
- List $/1M
- $2/10
14 Claude Opus 4.7*
- Hermes Index
- 71.20
- Index band
- 70.41–71.20
- Coverage
- 4/6 direct
- Value (Index÷task$)
- 2.37
- LMSYS
- 1498 (z=+1.11, w·z=+0.123) · as of 2026-09-25
- Agent Arena
- — (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
- Terminal-Bench 4.0
- — (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
- BrowseComp
- 79.3% (z=-0.47, w·z=-0.079) · as of 2026-09-25
- HLE
- 42.3% (z=-0.54, w·z=-0.089) · as of 2026-09-25
- GPQA Diamond
- 91.4% ±3.9 (z=-0.63, w·z=-0.035) · as of 2026-09-25
- SWE-Atlas-QnA
- — (absent; imputed z=0)
- DeepSWE
- — (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
- DeepSWE $/task
- —
- List $/1M
- $5/25
15 Claude Opus 4.6*
- Hermes Index
- 71.09
- Index band
- 70.17–71.09
- Coverage
- 4/6 direct
- Value (Index÷task$)
- 2.37
- LMSYS
- 1501 (z=+1.24, w·z=+0.138) · as of 2026-09-25
- Agent Arena
- — (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
- Terminal-Bench 4.0
- — (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
- BrowseComp
- 83.7% (z=+0.04, w·z=+0.007) · as of 2026-09-25
- HLE
- 39.9% (z=-0.92, w·z=-0.153) · as of 2026-09-25
- GPQA Diamond
- 89.6% ±4.2 (z=-1.49, w·z=-0.083) · as of 2026-09-25
- SWE-Atlas-QnA
- — (absent; imputed z=0)
- DeepSWE
- — (absent; imputed z=0) (z=+0.00, w·z=+0.000) · as of 2026-09-25
- DeepSWE $/task
- —
- List $/1M
- $5/25
* Fewest active benchmarks on this board.
Method
How this board is built.
Daily deterministic collectors. Meta boards are diagnostics only.
Composite formula
Scheme Agent/system 70% · Reasoning 20% · Chat 10%. AA · DeepSWE 20% each; TB · BrowseComp · HLE 15% each.
Composite score pipeline
Each active source contributes one metric. Percentage benches that arrive as 0–1 fractions are scaled to percent; there is no fixed bound map before ranking. For every source, we compute a robust center and scale on the fixed scored-model set (models with at least four active direct scores): the center is the median, and the scale is max(1.4826×MAD, ε), with ε = 1.0 percentage point for percent benches and ε = 10 Elo for LMSYS. Each model’s per-source z is (metric − center) / scale, then clamped to ±3. The Hermes Index is the weighted mean of those z values under the active weights for this run, mapped as clamp(72 + 10·z̄, 10, 100). It is an index, not a percentage. The primary path treats a missing active source as z = 0 (shrinkage to the board center). The sensitivity path drops missing sources and renormalizes the remaining weights; that peer is shown as a band only when coverage is incomplete. Sources that cover less than 80% of the scored set have their weight capped at 15% before renormalization so sparse boards cannot dominate. Stale or coverage-failed sources are dropped from the active set entirely. SWE-Atlas-QnA is not weighted (Hermes task mix is terminal/agent execution, not codebase QnA). BFCL-class and tau-bench were evaluated on 2026-08-30 against the freshness and frontier-hit floor and are absent from scoring.
| Source | Bucket | Nominal | Active | Effective | Floor |
|---|---|---|---|---|---|
| Terminal-Bench 4.0Agent/system | Agent/system | 15% | 16.7% | 11.4% | Direct |
| DeepSWEAgent/system | Agent/system | 20% | 16.7% | 17.4% | Direct |
| Agent ArenaAgent/system | Agent/system | 20% | 16.7% | 11.3% | Derived |
| BrowseCompAgent/system | Agent/system | 15% | 16.7% | 20.2% | Direct |
| HLEReasoning | Reasoning | 15% | 16.7% | 26.9% | Direct |
| LMSYS Text ArenaChat | Chat | 10% | 11.1% | 9.9% | Direct |
| GPQA DiamondReasoning | Reasoning | 5% | 5.6% | 2.9% | Direct |
Active = coverage-capped then renormalized. Effective = share of realized variance of per-source contributions across the scored board.
Agent Arena Agent Arena is a rank-linear percentile, 100×(N−rank)/(N−1), so rank 1 is exactly 100% and the board ends near 0%. Against the scored set that compresses z to roughly ±1.35 by construction — much tighter than free metric scales. It carries weight but does not count toward the ≥4 direct-source floor.
Coverage & missing A dash means the model has no score on that active source. The primary Index uses z = 0 for that cell; the sensitivity band shows weight renormalization and appears only when coverage is incomplete. Imputed cells do not count toward the inclusion floor.
Sources Tool-use board audit 2026-08-30: BFCL v4 (BenchLM) was ≤7d fresh with 13 rows but 0 models in the Hermes frontier seed after name cleaning — fails ≥5 frontier hits. tau-bench (BenchLM) published an empty leaderboard (display-only / outdated tasks). Neither is a primary source. SWE-Atlas-QnA weight is 0 (not Hermes task mix). Terminal-Bench is pinned to release 4.0 (BenchLM mirror of tbench.ai; best harness/effort per model). Scores are not comparable to the prior TB 2.x column; cross-version fallback into one z-pool is forbidden.
CIs Only GPQA currently shows a ±95% Wald interval (n≈198 multiple-choice items). LMSYS is Elo, not a binomial accuracy; Agent Arena is a rank transform; BrowseComp, Terminal-Bench 4.0, DeepSWE, and HLE are publisher point estimates without a stable public n for the same binomial model. Extending CIs needs per-source n or SEs.
MAD ε Scale floor agent_lmsys ε=1, browsecomp ε=1, deepswe ε=1, gpqa_diamond ε=1, hle ε=1, lmsys ε=10, terminalbench ε=1. Z clamped to ±3.0.
Picks Picks use the same top-15 board as the public leaderboard (models with ≥4 active directs, ranked by Hermes Index). The Hermes Index itself is capability-only and never uses price. Price enters only Cheap and Default. List $/1M means input+output USD per 1M tokens. Task $ means DeepSWE average $/task (real multi-step agent spend on that board). Frontier is the best Hermes Index (#1); the runner-up is the closest peer among ranks 2–5 within 2.0 Index points (higher Index then board rank on ties; no price), otherwise #2 by Index rank. Cheap is the lowest DeepSWE $/task inside that top-15 (list $/1M is only a tie-break or fallback when task $ is missing); its runner-up is the highest Index within 2× that task $ (else the next-cheapest by the same key). Default is the highest Hermes Index among models whose task $ is within 2× Cheap's task $ (list $ defines the band only when Cheap lacks task $), excluding Frontier #1 — the best daily driver still near the cheap frontier, not max Index÷$. The Default runner-up is the next-highest Index in that same band (if the band is a singleton, the next-cheapest model). Default may equal Cheap when Cheap also leads the band on Index; the rule never elevates a dominated model just to keep three category names distinct. Aggregate ≠ task proof.
Inclusion rules
- Hermes-oriented composite for personal agents — not a neutral AGI ranking.
- ≥4 active direct sources required (Agent Arena does not count toward the floor).
- List $/1M = input+output USD per 1M tokens. DeepSWE $ = avg $/task at best-effort config.
- * Fewest active benchmarks on this board.
- Aggregate ≠ task proof.
Source status
- LMSYS Text Arena — active, as of 2026-09-25, frontier hits 15
- HLE — active, as of 2026-09-25, frontier hits 16
- GPQA Diamond — active, as of 2026-09-25, frontier hits 16
- Terminal-Bench 4.0 — active, as of 2026-09-25, frontier hits 10
- BrowseComp — active, as of 2026-09-25, frontier hits 7
- DeepSWE — active, as of 2026-09-25, frontier hits 14
- Agent Arena — active, as of 2026-09-25, frontier hits 17