2024年全球竞赛推动通用智能评测得分从33%升至55.5%。
ARC Prize 2024: Technical Report
- 融合深度学习与程序合成,结合测试时训练提升推理能力。
- 顶尖模型在ARC-AGI私有集上得分达55.5%,较此前提升22.5个百分点。
- 适合关注通用人工智能与少样本泛化研究的学者和开发者。
截至2024年12月,ARC-AGI基准已存在五年,至今未被超越。我们认为它是当前世界最重要的未解人工智能基准,因其聚焦于对新任务的泛化能力——智能的本质,而非可预先准备的任务技能。今年,我们发起ARC Prize全球竞赛,旨在激发新思路,推动开放进展,目标是达到85%的基准分数。结果,ARC-AGI私有评估集上的最先进水平从33%提升至55.5%,得益于深度学习引导的程序合成与测试时训练等前沿通用智能推理技术。本文综述了顶尖方法,回顾了开源实现,讨论了ARC-AGI-1数据集的局限性,并分享了竞赛中的关键洞见。
原文摘要 · Abstract (English)
As of December 2024, the ARC-AGI benchmark is five years old and remains unbeaten. We believe it is currently the most important unsolved AI benchmark in the world because it seeks to measure generalization on novel tasks -- the essence of intelligence -- as opposed to skill at tasks that can be prepared for in advance. This year, we launched ARC Prize, a global competition to inspire new ideas and drive open progress towards AGI by reaching a target benchmark score of 85\%. As a result, the state-of-the-art score on the ARC-AGI private evaluation set increased from 33\% to 55.5\%, propelled by several frontier AGI reasoning techniques including deep learning-guided program synthesis and test-time training. In this paper, we survey top approaches, review new open-source implementations, discuss the limitations of the ARC-AGI-1 dataset, and share key insights gained from the competition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。