arXiv:2505.11831cs.AI2025-05被引 167

升级版推理基准,挑战AI更高阶抽象思维能力

ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems

  • 保留输入输出对格式,新增更精细的任务集
  • 人类可解但当前AI仍难应对,凸显认知复杂度
  • 适合评估前沿AI系统在类人智能上的进展

2019年推出的用于评估人工智能通用流体智力的抽象与推理语料库(ARC-AGI),通过一组新颖且无需先验知识的任务,成为极具挑战性的基准。过去五年间,它推动了大量研究。然而,随着AI技术进步,亟需更精细的评估工具以衡量更高层次的认知复杂性。本文提出ARC-AGI-2,作为该基准的升级版本。它延续原有输入输出对任务形式,确保研究连续性;同时引入全新策划与扩展的任务集,旨在提供更细粒度信号,以评估抽象推理与问题解决能力在更高水平流体智力下的表现。为刻画其难度与特性,我们报告了大规模人类测试结果,构建了可靠基线——表明该任务对人类可解,却对当前AI系统极具挑战。ARC-AGI-2旨在成为下一代工具,严格衡量向更通用、类人智能迈进的进展。

原文摘要 · Abstract (English)

The Abstraction and Reasoning Corpus for Artificial General Intelligence (ARC-AGI), introduced in 2019, established a challenging benchmark for evaluating the general fluid intelligence of artificial systems via a set of unique, novel tasks only requiring minimal prior knowledge. While ARC-AGI has spurred significant research activity over the past five years, recent AI progress calls for benchmarks capable of finer-grained evaluation at higher levels of cognitive complexity. We introduce ARC-AGI-2, an upgraded version of the benchmark. ARC-AGI-2 preserves the input-output pair task format of its predecessor, ensuring continuity for researchers. It incorporates a newly curated and expanded set of tasks specifically designed to provide a more granular signal to assess abstract reasoning and problem-solving abilities at higher levels of fluid intelligence. To contextualize the difficulty and characteristics of ARC-AGI-2, we present extensive results from human testing, providing a robust baseline that highlights the benchmark's accessibility to human intelligence, yet difficulty for current AI systems. ARC-AGI-2 aims to serve as a next-generation tool for rigorously measuring progress towards more general and human-like AI capabilities.

推理能力基准测试通用智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。