Skip to content

ARC-AGI Leaderboard

Technology

ARC-AGI has evolved from its first versions (ARC-AGI-1 and 2) which measured passive fluid intelligence, to ARC-AGI-3 which challenges AI agents to adapt on the fly to novel interactive environments.

The scatter plot above visualizes the critical relationship between cost-per-task and performance – a key measure of efficiency. True intelligence isn't just about solving problems, but solving them efficiently with minimal resources.

For more information, see our testing policy.

Only systems which required less than $10,000 to run are shown.

For models that were not able to produce full test out puts, remaining tasks were marked as incorrect.

Results marked as "preview" are unofficial and may be based on incomplete testing.

1 ARC-AGI-2 score estimate based on partial testing results and o1-pro pricing.

2 Provisional cost estimates based on Gemini 3 Pro pricing. Model to be retested once released.


Source: Hacker News — This article was automatically imported from the source. Read full article at original source →

HA
Originally published by Hacker News arcprize.org
Visit original article

Gram Slattery

Leave a Comment