arXiv:2501.07458cs.AIcs.PF2025-01被引 24

o3高分靠大量试错,不等于真正智能。

Understanding and Benchmarking Artificial Intelligence: OpenAI's o3 Is Not AGI

  • 用海量计算试错预设操作组合,是o3得分关键
  • 在特定任务上得87.5%,但无法应对真实世界未知问题
  • 提出新评测标准,推动真正通用智能研究

OpenAI的o3在ARC-AGI基准上取得87.5%的高分,引发对大语言模型是否具备智能及向通用人工智能(AGI)进展的思考。基于ARC-AGI提出者François Chollet关于技能与智能的区分,本文提出新智能定义:智能程度取决于代理在多样环境中以更少知识高效达成更多目标的能力。分析表明,ARC-AGI任务本质上是可通过大量试错预设操作组合解决的特定类型问题,o3正是通过大规模计算实现此策略并获得高分。然而,在物理世界与人类领域中,多数问题无法预先测试,且无现成操作可选,因此依赖预设操作的试错方法无法支撑真正的AGI。为此,本文提出新基准,覆盖更广泛的未知任务,以全面评估智能水平与向AGI的进展。

原文摘要 · Abstract (English)

OpenAI's o3 achieves a high score of 87.5 % on ARC-AGI, a benchmark proposed to measure intelligence. This raises the question whether systems based on Large Language Models (LLMs), particularly o3, demonstrate intelligence and progress towards artificial general intelligence (AGI). Building on the distinction between skills and intelligence made by François Chollet, the creator of ARC-AGI, a new understanding of intelligence is introduced: an agent is the more intelligent, the more efficiently it can achieve the more diverse goals in the more diverse worlds with the less knowledge. An analysis of the ARC-AGI benchmark shows that its tasks represent a very specific type of problem that can be solved by massive trialling of combinations of predefined operations. This method is also applied by o3, achieving its high score through the extensive use of computing power. However, for most problems in the physical world and in the human domain, solutions cannot be tested in advance and predefined operations are not available. Consequently, massive trialling of predefined operations, as o3 does, cannot be a basis for AGI - instead, new approaches are required that can reliably solve a wide variety of problems without existing skills. To support this development, a new benchmark for intelligence is outlined that covers a much higher diversity of unknown tasks to be solved, thus enabling a comprehensive assessment of intelligence and of progress towards AGI.

通用智能评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。