新基准测试挑战AI在抽象环境中的自主推理与规划能力
ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence
- 设计交互式抽象环境,要求智能体自主探索和建模
- 当前顶级AI系统在该任务上得分低于1%,人类可达100%
- 专为评估智能体的适应性效率而设,不依赖语言与外部知识
我们提出ARC-AGI-3,一个通过新颖、抽象、回合制环境研究代理智能的交互式基准。智能体需在无显式指令下自主探索、推断目标、构建环境动态的内部模型并规划有效动作序列。与前代ARC-AGI-1和2一致,本基准完全聚焦于评估智能体在新任务上的流体适应效率,避免使用语言和外部知识。环境仅基于核心知识先验,并通过大量人类测试进行难度校准。测试显示人类可解决100%的环境,而截至2026年3月,前沿AI系统得分低于1%。本文介绍了基准设计、基于人类动作基线的效率评分框架,以及环境构建、验证与校准的方法。
原文摘要 · Abstract (English)
We introduce ARC-AGI-3, an interactive benchmark for studying agentic intelligence through novel, abstract, turn-based environments in which agents must explore, infer goals, build internal models of environment dynamics, and plan effective action sequences without explicit instructions. Like its predecessors ARC-AGI-1 and 2, ARC-AGI-3 focuses entirely on evaluating fluid adaptive efficiency on novel tasks, while avoiding language and external knowledge. ARC-AGI-3 environments only leverage Core Knowledge priors and are difficulty-calibrated via extensive testing with human test-takers. Our testing shows humans can solve 100% of the environments, in contrast to frontier AI systems which, as of March 2026, score below 1%. In this paper, we present the benchmark design, its efficiency-based scoring framework grounded in human action baselines, and the methodology used to construct, validate, and calibrate the environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。