🏆 Scale AI official board · fetched 2026-10-07

SWE-Bench Pro 2026 — The Realistic Benchmark for AI Coding Models

As of 2026-10-07, Scale AI's official SWE-Bench Pro board (Public Dataset) lists Muse Spark 1.1* at 61.50 ± 3.10 and gpt-5.4 (xHigh)* at 59.10 ± 3.56, both Rank (UB) 1; on the Private Dataset the top two are Muse Spark 1.1* (51.5) and claude-opus-4-6 (thinking)* (47.1). For entries marked *, the board notes: Run with mini-swe-agent harness. Vendor self-reports and aggregated boards use other methodologies; see our 'Three First Places' page.

SWE-Bench Pro vs Verified — The Benchmark

SWE-Bench Verified (popular from 2024): ~500 human-verified GitHub issue-fix tasks. SWE-Bench Pro (mainstream from 2026): same approach but harder — longer context, more files touched, closer to real PR workflows. Verified plateaued near 80% (5.2-Codex), Pro still has ~40 points of headroom and is now the canonical benchmark. Terminal-Bench 2.0 measures terminal-agent tasks; OSWorld measures GUI-operating tasks; GDPval measures professional knowledge.

2026 Q1-Q2 Score Evolution

Early evolution: GPT-5.2-Codex (2026-01) Verified 80.0% / Pro 56.4%; GPT-5.3-Codex (2026-02) Pro 56.8% (this model was retired in 2026-08); GPT-5.5 Pro 58.6%. The 08-18 aggregated board (llm-stats / benchlm methodology): Mythos 5 80.3, Fable 5 80.0, Opus 5 79.2, Qwen3.8 Max 67.7, GPT-5.6 Sol 64.6, GLM-5.2 62.1. Note: Anthropic's system card self-reports Fable 5 at 80.3 (tied with Mythos), while the aggregated board records 80.0 — both numbers exist; the difference is the evaluation methodology.

Scale AI's Official Board, Verbatim: Public and Private Sets (2026-10-07)

Below is Scale AI's official SWE-Bench Pro board as fetched on 2026-10-07, copied without conversion and without model picks. Public Dataset (the board states 731 instances), top 6: Muse Spark 1.1* 61.50 (marked NEW), gpt-5.4 (xHigh)* 59.10, Muse Spark* 55.00, claude-opus-4-6 (thinking)* 51.90, gemini-3.1-pro (thinking)* 46.10, claude-opus-4-5-20251101 45.89. Private Dataset (276 instances sourced from 18 private, proprietary codebases from startups), top 4: Muse Spark 1.1* 51.5, claude-opus-4-6 (thinking)* 47.1, Muse Spark* 44.7, gpt-5.4(xHigh)* 43.4. Legend, verbatim: Rank (UB) is 1 + the number of models whose lower CI bound exceeds this model’s upper CI bound; * means Run with mini-swe-agent harness; entries that are not grayed out were run with an uncapped cost and with a turn limit of 250.

Unified Access to All Major Models via QCode

QCode.cc provides a single transparent API gateway to all major coding models from inside China — Claude (Opus 4.7 / Sonnet), GPT (5.5 / Codex family), Chinese models and more. One subscription, usage-based billing. In Claude Code, Codex CLI, Cursor, Cline, Continue, switching models is just changing base URL and model id — configure once, run everywhere.

FAQ

Is 58.6% on SWE-Bench Pro high? How should I read it?

Pro is far harder than Verified: Verified is saturated (80%+), while on Pro the top vendor self-reports sit around 80 and Scale's official unified-framework numbers come in lower. Before reading any Pro score, ask three things: which board, what scaffold, which effort tier — details in 'SWE-bench Pro's Three First Places'.

Why does general-purpose GPT-5.5 beat coding-specialized 5.3-Codex on Pro?

That's a first-half-of-2026 picture. The top of the 08-18 aggregated board is now three Claude models (Mythos 5 / Fable 5 / Opus 5), and GPT-5.3-Codex was retired in 2026-08 (migrate to the three GPT-5.6 tiers). Since 2026-09-03 the top of OpenAI's line is GPT-6 Astra — but note that OpenAI published no SWE-bench Verified score for it on either the model page or the system card, so it has no comparable position on this board. Do not treat numbers circulating elsewhere as official.

Opus 4.7's Pro score isn't published — how do I evaluate it?

On this 2026-08-18 aggregated board Opus 4.7 has no vendor-reported Pro score, so don't rank it by score. Its slot moved on to Opus 5; as of 2026-09-29 the newest Opus from Anthropic is Opus 5.5 (released 2026-09-22, $4 / $20). See the tracker /claude-opus-5-2-status-tracker and the specs in /claude-opus-5-5-guide.

Access All Major Coding Models Through QCode

claude-opus-5-5, claude-opus-5, claude-fable-5, gpt-6.1-sol, gpt-5.6-sol, glm-5.3, kimi-k3 and deepseek-v4-pro are callable with one QCode key. Billed per token; per-model prices are on /models.

Start Your QCode Plan

Create a QCode account

One key across Claude, GPT and Chinese models, billed per token — live rates on /models.

Try first, then decide

Not sure which tier? Start with Starter ($8.57/mo) and upgrade when you're happy — the unused value of the old plan goes back to your balance.