Benchmarks and leaderboards
SWE-bench, Terminal-Bench and other public results — and what they actually tell you.
SWE-Bench Pro ranking as of 2026-10-07: Scale AI's official public and private dataset boards, verbatim
As of 2026-10-07, Scale AI's official SWE-Bench Pro public board lists Muse Spark 1.1* at 61.50 and gpt-5.4 (xHigh)* at 59.10. Board text only, no model picks.
Updated 2026-10-08
Evaluating Ornith-1.5 yourself
Judge a new model against your own work instead of public leaderboards.
Updated 2026-10-08
AA Index Explained: v4.3.2 Weights and 10 Evals
The 10 evaluations and category weights as of 2026-10-07, compared with the old v4.1 weights
Updated 2026-10-07
Terminal-Bench 3.0 Explained
TB 3.0 vs. 2.1, current rankings, and which leaders are sellable
Updated 2026-09-21
Terminal-Bench 2.1 Leaderboard Explained
Three boards, three winners: the official 83.8 / AA 89.5 / vals 85.77 gap fully explained
Updated 2026-09-21
SWE-bench Pro's Three First Places Explained
80.3 / 80.0 / the official board: the truth about data splits and scaffolds
Updated 2026-08-22
From SWE-Bench to Production: How Benchmarks Do (Not) Translate to Delivery Speed in 2026
88% on SWE-Bench sounds great, but often only 30% on your private repos. Understand the gap and build your own verification harness.
Updated 2026-08-17
Claude Fable 5 SWE-Bench Pro Score, Explained
Where the 80.3% figure comes from, official methodology vs. independent reproduction, and how much it should weigh in your decision
Updated 2026-08-15