AI在抽象推理任务上进步缓慢,跨版本性能下降2-3倍,仍远不如人类。
The ARC of Progress towards AGI: A Living Survey of Abstraction and Reasoning
- 对比三版基准测试,程序合成、神经符号与纯神经方法均表现退化2-3倍
- 模型在最新版任务上仅达13%准确率,人类保持接近完美
- 低成本高效模型证明智能本质是技能获取效率,非单纯规模
抽象与推理语料库(ARC-AGI)已成为衡量AI流体智能的关键基准。本综述首次对82种方法在三个版本基准及ARC Prize 2024-2025竞赛中进行跨代分析。核心发现:所有范式——程序合成、神经符号与神经方法——在版本间均出现2-3倍性能下降,表明组合泛化存在根本性局限。当前系统在ARC-AGI-1上已达93.0%(Opus 4.6),但在ARC-AGI-2降至68.8%,在ARC-AGI-3仅为13%,而人类在所有版本中保持近似完美。成本一年内下降390倍(o3的$4,500/任务至GPT-5.2的$12/任务),但主要反映测试时并行度降低。万亿级模型得分与成本差异大,而受Kaggle约束的模型(660M-8B)表现竞争力,支持Chollet观点:智能即技能获取效率。测试时自适应与迭代优化成为关键成功因素,组合推理与交互学习仍无解。2025年冠军需数十万合成样本才在ARC-AGI-2上达24%,证实推理仍依赖知识。本版《ARC-AGI生活式综述》截至2026年2月,更新地址为https://nimi-ai.com/arc-survey/
原文摘要 · Abstract (English)
The Abstraction and Reasoning Corpus (ARC-AGI) has become a key benchmark for fluid intelligence in AI. This survey presents the first cross-generation analysis of 82 approaches across three benchmark versions and the ARC Prize 2024-2025 competitions. Our central finding is that performance degradation across versions is consistent across all paradigms: program synthesis, neuro-symbolic, and neural approaches all exhibit 2-3x drops from ARC-AGI-1 to ARC-AGI-2, indicating fundamental limitations in compositional generalization. While systems now reach 93.0% on ARC-AGI-1 (Opus 4.6), performance falls to 68.8% on ARC-AGI-2 and 13% on ARC-AGI-3, as humans maintain near-perfect accuracy across all versions. Cost fell 390x in one year (o3's $4,500/task to GPT-5.2's $12/task), although this largely reflects reduced test-time parallelism. Trillion-scale models vary widely in score and cost, while Kaggle-constrained entries (660M-8B) achieve competitive results, aligning with Chollet's thesis that intelligence is skill-acquisition efficiency. Test-time adaptation and refinement loops emerge as critical success factors, while compositional reasoning and interactive learning remain unsolved. ARC Prize 2025 winners needed hundreds of thousands of synthetic examples to reach 24% on ARC-AGI-2, confirming that reasoning remains knowledge-bound. This first release of the ARC-AGI Living Survey captures the field as of February 2026, with updates at https://nimi-ai.com/arc-survey/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。