Benchmark Explainer · AA Intelligence Index v4.3.2 · checked 2026-10-07

How the Artificial Analysis Intelligence Index Is Calculated
v4.3.2: a weighted average of 10 evaluations, with published weights

As of 2026-10-07, the AA Index (Artificial Analysis Intelligence Index) is on v4.3.2: a weighted average of 10 evaluations, with Agents 30%, Coding 20%, Scientific Reasoning 20% and General 30%. The widely quoted 34% / 24% / 24% / 18% split is the older v4.1 weighting from June 2026; when a model's score differs between sites, index version or data-pull timing is the usual cause.

Updated 2026-10-07

Beyond the leaderboard — one key to benchmark Claude, GPT and Chinese models yourself.

#AA Intelligence Index#v4.3.2#10 weighted evals#pass@1

Four key facts

30%

Agents category

AA-Briefcase v1.1 at 15%, GDPval-AA v2.1 at 10%, AutomationBench-AA at 5%; Artificial Analysis says the weighting emphasizes agentic tasks.

20% + 20%

Coding and Scientific Reasoning

Coding is 20% (Terminal-Bench 4.0 10%, SciCode 10%); Scientific Reasoning is 20% (Humanity's Last Exam 10%, CritPt 10%). GPQA Diamond left the index in v4.2.

30%

General category

AA-Omniscience 15% (10% accuracy plus 5% non-hallucination), GDP.pdf 10%, AA-LCR v1.1 5%. Under v4.1, General was only 18%.

pass@1

Scoring method

Terminal-Bench 4.0, SciCode, AA-LCR, HLE and CritPt are scored pass@1; the Agents evaluations mainly use Elo from judge-panel pairwise comparisons.

The 10 evaluations in v4.3.2 and their weights

As of 2026-10-07, the current AA Intelligence Index, v4.3.2, is a weighted average of 10 evaluations in four categories: Agents 30% (AA-Briefcase v1.1 15%, GDPval-AA v2.1 10%, AutomationBench-AA 5%); Coding 20% (Terminal-Bench 4.0 10%, SciCode 10%); General 30% (AA-Omniscience 15%, split into 10% accuracy and 5% non-hallucination; GDP.pdf 10%; AA-LCR v1.1 5%); Scientific Reasoning 20% (Humanity's Last Exam 10%, CritPt 10%). Artificial Analysis says the weighting emphasizes agentic tasks. Scoring is not uniform: Terminal-Bench 4.0, SciCode, AA-LCR, HLE and CritPt use pass@1; AA-Briefcase and GDPval-AA use Elo from pairwise comparisons; AutomationBench-AA scores objective completion.

Why the same model's score looks different in different places

Two models with the same composite score don't necessarily have the same capability profile — different strengths can sum to the same weighted total. Artificial Analysis also keeps revising the index: v4.1 (June 2026) weighted Agents 34%, Coding 24%, Scientific Reasoning 24% and General 18%; September 2026 brought v4.2, v4.3, v4.3.1 and v4.3.2, and as of 2026-10-07 the current version is v4.3.2 with 10 evaluations. Versions differ in evaluations and weights; even within one version, a different data-pull date can capture model updates, so the same model's score can drift.

Timeline

Early versions

Artificial Analysis introduces the Intelligence Index as a weighted composite score for comparing model capability across the board.

2026-06

Index v4.1 is introduced, publicly specifying the nine sub-evaluations and their weights (Agents 34%, Coding 24%, Scientific Reasoning 24%, General 18%).

2026-09

September 2026: v4.2, v4.3, v4.3.1 and v4.3.2 ship. v4.2 adds AA-Briefcase (15%) and GDP.pdf (10%) and removes GPQA Diamond; v4.3 replaces τ³-Banking with AutomationBench-AA and Terminal-Bench 2.1 with Terminal-Bench 4.0. Current category weights: Agents 30%, Coding 20%, Scientific Reasoning 20%, General 30%.

Confirmed vs. common misreading

Confirmed

As of 2026-10-07, the official methodology page states: v4.3.2 incorporates 10 evaluations; category weights are Agents 30%, Coding 20%, Scientific Reasoning 20%, General 30%; per-evaluation weights are AA-Briefcase 15%, GDPval-AA 10%, AutomationBench-AA 5%, Terminal-Bench 4.0 10%, SciCode 10%, AA-LCR 5%, AA-Omniscience accuracy 10% and non-hallucination 5%, HLE 10%, GDP.pdf 10%, CritPt 10%. Its version history lists v4.1 (June to August 2026) at 34% / 24% / 24% / 18%. Artificial Analysis estimates the index's 95% confidence interval at less than ±1%.

Common misreading

Many people treat the composite score as a single-dimension ranking and use it to claim 'this one beats that one' — but it is a weighted sum of ten heterogeneous evaluations, so two models with similar totals can have very different profiles across Agents, Coding, Scientific Reasoning and General. Another common misreading is explaining a v4.3.2 score with the v4.1 weights (34% / 24% / 24% / 18%); the two weightings are not the same.

Composite score vs. sub-scores

Just the composite score

Good for a quick rough tier of capability, but doesn't tell you specifically where a model is strong or weak.

Breaking it down by sub-evaluation

If your use case leans heavily toward Agents or Coding, going straight to the relevant sub-evaluation (such as Terminal-Bench 4.0 in the v4.3.2 Coding category) is more targeted than the composite.

How to actually use this index

Step 1: figure out which of the four categories (Agents, Coding, Scientific Reasoning or General) is closest to your use case, and prioritize the relevant sub-score over the composite. Step 2: when comparing models, confirm both scores come from the same index version (v4.3.2 as of 2026-10-07) and the same data pull — cross-version or cross-date comparisons are easy to misread. Step 3: treat the composite as a rough filter, not a final answer — pair it with a small-scale test on your actual task before deciding.

What to do on QCode

As of 2026-10-07, one QCode key can call claude-opus-5-5, claude-sonnet-5-5, gpt-6.1-sol, deepseek-v4-pro, glm-5.3, qwen3.8-max, kimi-k3 and other models, so you can run a small comparison on your own task instead of deciding off one composite number. Billing is per token; per-model prices are on /models, and model availability follows /models.

FAQ

How exactly is the AA Intelligence Index calculated?

As of 2026-10-07, v4.3.2 is a weighted average of 10 evaluations: Agents 30% (AA-Briefcase 15% + GDPval-AA 10% + AutomationBench-AA 5%), Coding 20% (Terminal-Bench 4.0 10% + SciCode 10%), General 30% (AA-Omniscience accuracy 10% + non-hallucination 5% + GDP.pdf 10% + AA-LCR 5%), Scientific Reasoning 20% (HLE 10% + CritPt 10%).

Why does the same model's score look different than what I saw before?

The most common reasons are an index version upgrade (weights or sub-evaluation composition changed) or a different data-pull date (the model itself may have been updated, or the eval environment adjusted).

Do two models with the same composite score really have equal capability?

Not necessarily. The composite is a weighted sum, so two models can each be strong and weak in different sub-areas and just happen to land on a similar total — their actual capability profiles can differ a lot.

What does pass@1 mean?

The model gets one shot and only passes if it is correct on the first attempt. In v4.3.2, Terminal-Bench 4.0, SciCode, AA-LCR, HLE and CritPt are scored pass@1; AA-Briefcase and GDPval-AA in the Agents category use pairwise Elo instead.

Terminal-Bench is only part of the Coding category — why do other pages on this site treat it as a standalone leaderboard?

Terminal-Bench is both an independent public leaderboard and a component of the AA Index's Coding category: v4.3 replaced Terminal-Bench 2.1 with Terminal-Bench 4.0, which carries 10% of the v4.3.2 total (Coding is 20%). Artificial Analysis notes that Terminal-Bench 2.1 remains part of its Coding Index. The two framings are just different levels of granularity.

What else should I watch for when picking a model based on this index?

Treat the index as a general reference, not the only input. If your task is specific, go straight to the relevant sub-evaluation's score, and pair it with a small-scale test before making a final call.

Sources

The ten v4.3.2 evaluations, category and per-evaluation weights, scoring methods, and the version history from v4.1 to v4.3.2 come from Artificial Analysis's official methodology page (Intelligence Benchmarking, including Version History) and the index notes on its homepage. Captured 2026-10-07. The historical v4.1 weights are taken from the same version history.

Don't decide off a single composite number

One QCode key calls Claude, GPT, DeepSeek, GLM, Qwen and Kimi models, so you can test them on your actual task.

Related reading

This page explains the Artificial Analysis Intelligence Index methodology based on its official public documentation, checked 2026-10-07; it is not an official QCode statement. Weights and evaluation composition follow Artificial Analysis's official page; model availability follows /models.

Try first, then decide

Not sure which tier? Start with Starter ($8.57/mo) and upgrade when you're happy — the unused value of the old plan goes back to your balance.