用几何框架重新定义AI评估,让通用智能进步看得见。
The Geometry of Benchmarks: A New Path Toward AGI
- 把评测集看作空间中的点,构建可度量的智能评估几何体系。
- 提出自进化系数κ,证明在特定条件下系统能自我提升。
- 适合关注通用智能演进路径的研究者与技术决策者。
评测是衡量人工智能进展的主要工具,但现有方法仅孤立评估模型在单个测试集上的表现,缺乏对泛化能力或自主自进化能力的指导。本文提出一种几何框架,将所有用于人工智能代理的心理测量测试集视为结构化模空间中的点,代理性能由该空间上的能力泛函描述。首先,定义了自主人工智能(AAI)等级,一个基于跨任务族(如推理、规划、工具使用和长程控制)测评表现的卡达舍夫式自主性层级。其次,构建评测集的模空间,识别在代理排序与能力推断层面不可区分的评测等价类;该几何结构带来确定性结果:稠密的评测族足以验证整个任务空间中的性能。第三,引入通用生成-验证-更新(GVU)算子,涵盖强化学习、自对弈、辩论和验证器微调等多种范式,并定义自进化系数κ为能力泛函沿诱导流的李导数。生成与验证联合噪声的方差不等式提供了κ > 0的充分条件。研究结果表明,通向人工通用智能(AGI)的进步应被理解为由GVU动态驱动的评测模空间上的流动,而非单一排行榜分数的提升。
原文摘要 · Abstract (English)
Benchmarks are the primary tool for assessing progress in artificial intelligence (AI), yet current practice evaluates models on isolated test suites and provides little guidance for reasoning about generality or autonomous self-improvement. Here we introduce a geometric framework in which all psychometric batteries for AI agents are treated as points in a structured moduli space, and agent performance is described by capability functionals over this space. First, we define an Autonomous AI (AAI) Scale, a Kardashev-style hierarchy of autonomy grounded in measurable performance on batteries spanning families of tasks (for example reasoning, planning, tool use and long-horizon control). Second, we construct a moduli space of batteries, identifying equivalence classes of benchmarks that are indistinguishable at the level of agent orderings and capability inferences. This geometry yields determinacy results: dense families of batteries suffice to certify performance on entire regions of task space. Third, we introduce a general Generator-Verifier-Updater (GVU) operator that subsumes reinforcement learning, self-play, debate and verifier-based fine-tuning as special cases, and we define a self-improvement coefficient $κ$ as the Lie derivative of a capability functional along the induced flow. A variance inequality on the combined noise of generation and verification provides sufficient conditions for $κ> 0$. Our results suggest that progress toward artificial general intelligence (AGI) is best understood as a flow on moduli of benchmarks, driven by GVU dynamics rather than by scores on individual leaderboards.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。