How the Artificial Analysis Intelligence Index Is Calculated
v4.3.2: a weighted average of 10 evaluations, with published weights
As of 2026-10-07, the AA Index (Artificial Analysis Intelligence Index) is on v4.3.2: a weighted average of 10 evaluations, with Agents 30%, Coding 20%, Scientific Reasoning 20% and General 30%. The widely quoted 34% / 24% / 24% / 18% split is the older v4.1 weighting from June 2026; when a model's score differs between sites, index version or data-pull timing is the usual cause.
Updated 2026-10-07
Beyond the leaderboard — one key to benchmark Claude, GPT and Chinese models yourself.
- Billed per token — live rates on /models
- Pay by card (Visa / Mastercard / AMEX), Apple Pay, Google Pay or crypto
- Self-service: top up and activate instantly
Four key facts
Agents category
AA-Briefcase v1.1 at 15%, GDPval-AA v2.1 at 10%, AutomationBench-AA at 5%; Artificial Analysis says the weighting emphasizes agentic tasks.
Coding and Scientific Reasoning
Coding is 20% (Terminal-Bench 4.0 10%, SciCode 10%); Scientific Reasoning is 20% (Humanity's Last Exam 10%, CritPt 10%). GPQA Diamond left the index in v4.2.
General category
AA-Omniscience 15% (10% accuracy plus 5% non-hallucination), GDP.pdf 10%, AA-LCR v1.1 5%. Under v4.1, General was only 18%.
Scoring method
Terminal-Bench 4.0, SciCode, AA-LCR, HLE and CritPt are scored pass@1; the Agents evaluations mainly use Elo from judge-panel pairwise comparisons.
The 10 evaluations in v4.3.2 and their weights
As of 2026-10-07, the current AA Intelligence Index, v4.3.2, is a weighted average of 10 evaluations in four categories: Agents 30% (AA-Briefcase v1.1 15%, GDPval-AA v2.1 10%, AutomationBench-AA 5%); Coding 20% (Terminal-Bench 4.0 10%, SciCode 10%); General 30% (AA-Omniscience 15%, split into 10% accuracy and 5% non-hallucination; GDP.pdf 10%; AA-LCR v1.1 5%); Scientific Reasoning 20% (Humanity's Last Exam 10%, CritPt 10%). Artificial Analysis says the weighting emphasizes agentic tasks. Scoring is not uniform: Terminal-Bench 4.0, SciCode, AA-LCR, HLE and CritPt use pass@1; AA-Briefcase and GDPval-AA use Elo from pairwise comparisons; AutomationBench-AA scores objective completion.
Why the same model's score looks different in different places
Two models with the same composite score don't necessarily have the same capability profile — different strengths can sum to the same weighted total. Artificial Analysis also keeps revising the index: v4.1 (June 2026) weighted Agents 34%, Coding 24%, Scientific Reasoning 24% and General 18%; September 2026 brought v4.2, v4.3, v4.3.1 and v4.3.2, and as of 2026-10-07 the current version is v4.3.2 with 10 evaluations. Versions differ in evaluations and weights; even within one version, a different data-pull date can capture model updates, so the same model's score can drift.
Timeline
Artificial Analysis introduces the Intelligence Index as a weighted composite score for comparing model capability across the board.
Index v4.1 is introduced, publicly specifying the nine sub-evaluations and their weights (Agents 34%, Coding 24%, Scientific Reasoning 24%, General 18%).
September 2026: v4.2, v4.3, v4.3.1 and v4.3.2 ship. v4.2 adds AA-Briefcase (15%) and GDP.pdf (10%) and removes GPQA Diamond; v4.3 replaces τ³-Banking with AutomationBench-AA and Terminal-Bench 2.1 with Terminal-Bench 4.0. Current category weights: Agents 30%, Coding 20%, Scientific Reasoning 20%, General 30%.
Confirmed vs. common misreading
Confirmed
As of 2026-10-07, the official methodology page states: v4.3.2 incorporates 10 evaluations; category weights are Agents 30%, Coding 20%, Scientific Reasoning 20%, General 30%; per-evaluation weights are AA-Briefcase 15%, GDPval-AA 10%, AutomationBench-AA 5%, Terminal-Bench 4.0 10%, SciCode 10%, AA-LCR 5%, AA-Omniscience accuracy 10% and non-hallucination 5%, HLE 10%, GDP.pdf 10%, CritPt 10%. Its version history lists v4.1 (June to August 2026) at 34% / 24% / 24% / 18%. Artificial Analysis estimates the index's 95% confidence interval at less than ±1%.
Common misreading
Many people treat the composite score as a single-dimension ranking and use it to claim 'this one beats that one' — but it is a weighted sum of ten heterogeneous evaluations, so two models with similar totals can have very different profiles across Agents, Coding, Scientific Reasoning and General. Another common misreading is explaining a v4.3.2 score with the v4.1 weights (34% / 24% / 24% / 18%); the two weightings are not the same.
Composite score vs. sub-scores
Just the composite score
Good for a quick rough tier of capability, but doesn't tell you specifically where a model is strong or weak.
Breaking it down by sub-evaluation
If your use case leans heavily toward Agents or Coding, going straight to the relevant sub-evaluation (such as Terminal-Bench 4.0 in the v4.3.2 Coding category) is more targeted than the composite.
How to actually use this index
Step 1: figure out which of the four categories (Agents, Coding, Scientific Reasoning or General) is closest to your use case, and prioritize the relevant sub-score over the composite. Step 2: when comparing models, confirm both scores come from the same index version (v4.3.2 as of 2026-10-07) and the same data pull — cross-version or cross-date comparisons are easy to misread. Step 3: treat the composite as a rough filter, not a final answer — pair it with a small-scale test on your actual task before deciding.
What to do on QCode
As of 2026-10-07, one QCode key can call claude-opus-5-5, claude-sonnet-5-5, gpt-6.1-sol, deepseek-v4-pro, glm-5.3, qwen3.8-max, kimi-k3 and other models, so you can run a small comparison on your own task instead of deciding off one composite number. Billing is per token; per-model prices are on /models, and model availability follows /models.
FAQ
How exactly is the AA Intelligence Index calculated?
As of 2026-10-07, v4.3.2 is a weighted average of 10 evaluations: Agents 30% (AA-Briefcase 15% + GDPval-AA 10% + AutomationBench-AA 5%), Coding 20% (Terminal-Bench 4.0 10% + SciCode 10%), General 30% (AA-Omniscience accuracy 10% + non-hallucination 5% + GDP.pdf 10% + AA-LCR 5%), Scientific Reasoning 20% (HLE 10% + CritPt 10%).
Why does the same model's score look different than what I saw before?
The most common reasons are an index version upgrade (weights or sub-evaluation composition changed) or a different data-pull date (the model itself may have been updated, or the eval environment adjusted).
Do two models with the same composite score really have equal capability?
Not necessarily. The composite is a weighted sum, so two models can each be strong and weak in different sub-areas and just happen to land on a similar total — their actual capability profiles can differ a lot.
What does pass@1 mean?
The model gets one shot and only passes if it is correct on the first attempt. In v4.3.2, Terminal-Bench 4.0, SciCode, AA-LCR, HLE and CritPt are scored pass@1; AA-Briefcase and GDPval-AA in the Agents category use pairwise Elo instead.
Terminal-Bench is only part of the Coding category — why do other pages on this site treat it as a standalone leaderboard?
Terminal-Bench is both an independent public leaderboard and a component of the AA Index's Coding category: v4.3 replaced Terminal-Bench 2.1 with Terminal-Bench 4.0, which carries 10% of the v4.3.2 total (Coding is 20%). Artificial Analysis notes that Terminal-Bench 2.1 remains part of its Coding Index. The two framings are just different levels of granularity.
What else should I watch for when picking a model based on this index?
Treat the index as a general reference, not the only input. If your task is specific, go straight to the relevant sub-evaluation's score, and pair it with a small-scale test before making a final call.
Sources
The ten v4.3.2 evaluations, category and per-evaluation weights, scoring methods, and the version history from v4.1 to v4.3.2 come from Artificial Analysis's official methodology page (Intelligence Benchmarking, including Version History) and the index notes on its homepage. Captured 2026-10-07. The historical v4.1 weights are taken from the same version history.
Don't decide off a single composite number
One QCode key calls Claude, GPT, DeepSeek, GLM, Qwen and Kimi models, so you can test them on your actual task.
Related reading
Terminal-Bench 3.0 explained
A detailed look at Terminal-Bench 3.0; from AA Index v4.3 the Coding category uses Terminal-Bench 4.0.
AI Model Radar 2026
Our own composite model comparison, useful to cross-reference against the AA Index.
SWE-bench Pro 2026 ranking
Another leading coding-focused leaderboard.
This page explains the Artificial Analysis Intelligence Index methodology based on its official public documentation, checked 2026-10-07; it is not an official QCode statement. Weights and evaluation composition follow Artificial Analysis's official page; model availability follows /models.