2025年ARC竞赛揭示了迭代优化对智能系统的重要性。
ARC Prize 2025: Technical Report
- 引入任务级迭代优化循环,提升模型泛化能力
- 顶尖模型在新数据集上最高达24%准确率
- 适合关注通用智能与推理机制的研究者
ARC-AGI基准系列是衡量少样本泛化能力的关键指标,体现了智能的核心特征。2025年全球竞赛聚焦于新发布的ARC-AGI-2数据集,其任务复杂度高于前代。该Kaggle竞赛吸引1,455支队伍提交15,154个方案,顶尖成绩在私有评估集上达到24%。论文投稿量接近翻倍至90篇,反映出对流体智能与抽象推理研究兴趣的持续增长。2025年的核心趋势是‘精炼循环’——一种由反馈信号驱动的、针对每个任务的迭代程序优化机制。该循环形式多样,包括进化式程序合成及商用AI系统的应用层优化。权重空间中的精炼也已实现,如零预训练深度学习方法仅用700万参数即取得竞争力表现。同时,四家前沿实验室(Anthropic、Google DeepMind、OpenAI、xAI)在2025年公开模型卡报告其在ARC-AGI上的表现,确立该基准为行业标准。然而分析表明,当前先进模型仍受限于知识覆盖范围,引发新型基准污染问题。本文综述顶级方法,探讨精炼循环在通用智能发展中的作用,分析知识依赖性过拟合,并预告将引入交互式推理挑战的ARC-AGI-3,要求探索、规划、记忆、目标获取与对齐能力。
原文摘要 · Abstract (English)
The ARC-AGI benchmark series serves as a critical measure of few-shot generalization on novel tasks, a core aspect of intelligence. The ARC Prize 2025 global competition targeted the newly released ARC-AGI-2 dataset, which features greater task complexity compared to its predecessor. The Kaggle competition attracted 1,455 teams and 15,154 entries, with the top score reaching 24% on the ARC-AGI-2 private evaluation set. Paper submissions nearly doubled year-over-year to 90 entries, reflecting the growing research interest in fluid intelligence and abstract reasoning. The defining theme of 2025 is the emergence of the refinement loop -- a per-task iterative program optimization loop guided by a feedback signal. Refinement loops come in a variety of forms, in particular evolutionary program synthesis approaches and application-layer refinements to commercial AI systems. Such refinement loops are also possible in weight space, as evidenced by zero-pretraining deep learning methods which are now achieving competitive performance with remarkably small networks (7M parameters). In parallel, four frontier AI labs (Anthropic, Google DeepMind, OpenAI, and xAI) reported ARC-AGI performance in public model cards in 2025, establishing ARC-AGI as an industry standard benchmark for AI reasoning. However, our analysis indicates that current frontier AI reasoning performance remains fundamentally constrained to knowledge coverage, giving rise to new forms of benchmark contamination. In this paper, we survey the top-performing methods, examine the role of refinement loops in AGI progress, discuss knowledge-dependent overfitting, and preview ARC-AGI-3, which introduces interactive reasoning challenges that require exploration, planning, memory, goal acquisition, and alignment capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。