arXiv:2510.01687cs.AIcs.CL2025-10

用数据科学方法评估AGI,强调实证能力而非主观任务设计。

Improving AGI Evaluation: A Data Science Perspective

  • 以真实任务执行能力为核心,替代依赖直觉的合成测试。
  • 借鉴数据科学中的可靠性验证流程,提升评估可信度。
  • 适合关注可落地AGI评测的科研与工程人员。

AGI系统的评估因目标范围广泛而困难,目前缺乏对最终状态的完美评估方法,只能通过小型测试来判断是否接近AGI。本文指出,当前评估方法多依赖对智能的主观直觉设计合成任务,历史表现不佳。我们主张采用新范式:聚焦系统在复杂任务中的稳健执行能力,以证明其具备通用智能。该思路源于数据科学中确保系统可部署的实践,强调结果可重复、性能稳定。文中提供具体应用案例,说明如何将此理念应用于实际评估体系中。

原文摘要 · Abstract (English)

Evaluation of potential AGI systems and methods is difficult due to the breadth of the engineering goal. We have no methods for perfect evaluation of the end state, and instead measure performance on small tests designed to provide directional indication that we are approaching AGI. In this work we argue that AGI evaluation methods have been dominated by a design philosophy that uses our intuitions of what intelligence is to create synthetic tasks, that have performed poorly in the history of AI. Instead we argue for an alternative design philosophy focused on evaluating robust task execution that seeks to demonstrate AGI through competence. This perspective is developed from common practices in data science that are used to show that a system can be reliably deployed. We provide practical examples of what this would mean for AGI evaluation.

AGI评估数据科学任务执行

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。