arXiv:2505.12844cs.AIcs.RO2025-05NeurIPS

用博弈评分法量化模型与任务难度,看清AI离全能还差多远。

AGI-Elo: How Far Are We From Mastering A Task?

  • 基于模型与任务的对抗评分,同时评估难度与能力
  • 跨视觉语言动作三领域验证,结果具可比性
  • 揭示真实任务长尾分布和当前模型差距

随着人工智能向通用智能(AGI)迈进,亟需超越平均性能指标的更全面评估体系。本文提出一种统一评分系统,联合建模个体测试题的难度与模型(或人类)在视觉、语言、行动领域的胜任力。不同于仅关注模型表现的现有方法,本方法通过模型与任务间的竞争互动,实现细粒度、难度感知的评估,捕捉现实挑战的长尾分布及当前模型与全任务掌控之间的能力差距。我们在多个成熟数据集和模型上开展广泛实验,验证了该系统的泛化性与鲁棒性。生成的评分分布为任务难度、模型演进路径以及通向全任务掌握仍存挑战提供了全新视角与可解释洞察。

原文摘要 · Abstract (English)

As the field progresses toward Artificial General Intelligence (AGI), there is a pressing need for more comprehensive and insightful evaluation frameworks that go beyond aggregate performance metrics. This paper introduces a unified rating system that jointly models the difficulty of individual test cases and the competency of AI models (or humans) across vision, language, and action domains. Unlike existing metrics that focus solely on models, our approach allows for fine-grained, difficulty-aware evaluations through competitive interactions between models and tasks, capturing both the long-tail distribution of real-world challenges and the competency gap between current models and full task mastery. We validate the generalizability and robustness of our system through extensive experiments on multiple established datasets and models across distinct AGI domains. The resulting rating distributions offer novel perspectives and interpretable insights into task difficulty, model progression, and the outstanding challenges that remain on the path to achieving full AGI task mastery.

AGI评估任务难度模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。