arXiv:2510.02125cs.AIcs.CL2025-10被引 12

AI模型在抽象推理上表现不如人类,视觉模态下能力被低估。

Do AI Models Perform Human-like Abstract Reasoning Across Modalities?

  • 用ConceptARC测试跨模态抽象能力,对比文本与视觉输入。
  • 最佳模型规则多依赖表面捷径,真正理解抽象的少于人类。
  • 视觉模态下准确率低但规则质量高,需结合解释分析评估。

OpenAI的o3-preview模型在ARC-AGI-1基准上超过人类准确率,但这是否意味着其真正具备所测抽象推理能力?我们使用更简单的ConceptARC基准评估了AI模型的抽象能力,考察输入模态(文本/视觉)、外部工具使用和推理投入。除了输出准确率,还分析模型生成的自然语言规则,以判断其是否识别出设计意图的抽象概念。结果表明,最佳模型的规则常基于表面‘捷径’,捕捉预期抽象的能力远低于人类。在视觉模态下,模型准确率显著下降;但规则层面分析显示,仍有相当比例的规则正确识别了目标抽象,即使未能正确应用。这说明仅凭准确率会高估文本模态下的能力,低估视觉模态下的潜力。研究提供了更真实的评估视角,推动迈向以抽象为核心的类人智能。

原文摘要 · Abstract (English)

OpenAI's o3-preview reasoning model exceeded human accuracy on the ARC-AGI-1 benchmark, but does that mean state-of-the-art models recognize and reason with the abstractions the benchmark was designed to test? Here we investigate abstraction abilities of AI models using the closely related but simpler ConceptARC benchmark. Our evaluations vary input modality (textual vs. visual), use of external Python tools, and reasoning effort. Beyond output accuracy, we evaluate the natural-language rules that models generate to explain their solutions, enabling us to assess whether models recognize the abstractions that ConceptARC was designed to elicit. We show that the best models' rules are frequently based on surface-level ``shortcuts,'' capturing intended abstractions considerably less often than humans. In the visual modality, AI models' output accuracy drops sharply; however, our rule-level analysis reveals that a substantial share of their rules capture the intended abstractions, even as the models struggle to apply these concepts to generate correct solutions. In short, we show that using accuracy alone to evaluate abstract reasoning can substantially overestimate AI capabilities in textual modalities and underestimate it in visual modalities. Our results offer a more faithful picture of AI models' abstract reasoning abilities and a more principled way to track progress toward human-like, abstraction-centered intelligence.

抽象推理多模态模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。