Artificial Analysis has released version 4.2 of its Intelligence Index, changing both the tasks and the balance of evidence used to rank frontier AI models. The update adds AA-Briefcase, a private agentic knowledge-work evaluation, and GDP.pdf, a long-context document-reasoning benchmark built around 4,592 PDF pages. It also removes GPQA Diamond after Artificial Analysis judged the benchmark saturated.\n\nThe larger methodological shift is toward held-out testing. Artificial Analysis says v4.2 puts greater weight on private test sets and upgrades its grading infrastructure, with the stated goal of making the benchmark harder to game and more representative of real work. Independent coverage reports that private held-out tests now account for 40% of the index weighting, double their share in v4.1.\n\nThat matters because benchmark contamination is becoming a practical problem for buyers as well as researchers. A score is less useful when the model developer may have seen the benchmark questions, or when top models have already saturated a test. Private task sets and newer real-world evaluations do not eliminate that problem, but they make a composite leaderboard less dependent on familiar public exams.\n\nThe refreshed leaderboard places Anthropic’s Claude Fable 5.1 at the top and OpenAI’s GPT-6 Astra close behind, while the new sub-tests show different leaders on different kinds of work. Those rankings should not be treated as a universal product verdict: the Index is still a composite benchmark, and AiToolMap uses benchmark results only when the tested model and task surface match the claim being made.\n\nFor AiToolMap, the main operational consequence is methodological. Existing rating evidence tied to Artificial Analysis remains useful, but any comparison that mixes pre-v4.2 and v4.2 Index scores needs to account for the changed benchmark composition rather than treating the numbers as perfectly continuous across versions.
← ALL NEWS
UPDATE · 2026-09-05
Artificial Analysis updates its Intelligence Index with tougher private tests
Intelligence Index v4.2 adds an agentic knowledge-work benchmark, a 4,592-page document-reasoning evaluation and more weight on held-out private tests. The methodological change matters as much as the new leaderboard because it is explicitly designed to reduce benchmark gaming and saturation.