arXiv:2607.00913cs.AI2026-07中稿 · ICML

不同评估方式下,大模型能力差距可能扩大或缩小,影响谁掌控AI技术。

Two AI Metrics Diverged: Will it Make All the Difference?

论文配图:Two AI Metrics Diverged: Will it Make All the Difference?
图 1 · 摘自论文原文
  • 按训练/推理算力划分指标函数形式,判断哪些指标利于小模型
  • 验证损失差距缩小,但其他指标显示大模型优势持续扩大
  • 关键能力若无界,将集中于少数巨头;若有界,则可普惠大众

随着算力指数级增长,前沿AI模型的能力是否会超越小预算开发者的可及范围?抑或能力趋于收敛,让‘弱小模型继承地球’?基于Gundlach等(2025b)的工作,我们发现答案取决于如何衡量和估值AI能力。分析传统性能指标后发现,尽管验证损失差距缩小,但在其他指标上,前沿模型的优势却持续扩大。通过将性能指标按其与训练(及推理)算力的函数关系分类,我们给出判断哪些指标有利于小模型的严格数学条件,并证明有界性能指标始终利于小模型。但需谨慎解读:许多常见有界指标存在密切相关的无界对应项(反之亦然)。在特定领域选择恰当指标是政策制定的前提,因为有界与无界指标可能导向相反政策建议。若某能力(如软件工程、合成生物学或修辞说服力)在我们关心的度量下为无界,则该能力将集中于少数富裕主体手中;反之,若有界,则前沿能力将通过小模型广泛传播至大众之手。

原文摘要 · Abstract (English)

As exponential compute scaling continues, will the capabilities of frontier AI models outstrip what is accessible to developers on a small fixed budget? Or will capabilities converge, with "meek models inheriting the earth"? Building on Gundlach et al. (2025b), we show that the answer depends on how we value and measure AI capabilities. We discuss conventional performance measures and show that, while validation loss shows a shrinking gap, on other metrics frontier models grow their lead forever. Classifying performance metrics by their functional forms in relation to training (and inference) compute, we provide tight mathematical conditions for determining which metrics favor meek models, and show that bounded performance metrics always do. But careful interpretation of performance metrics is essential: we show that many common bounded metrics have closely-related counterpart metrics that are unbounded (and vice versa). Determining the apt metric in a domain is a prerequisite for policy, since bounded and unbounded metrics may suggest opposing policy responses. If a particular capability -- like software engineering, synthetic biology, or rhetorical persuasiveness -- is unbounded when measured in the terms we care about, frontier-level capability will likely be concentrated in the hands of a few wealthy actors. Conversely, if that capability is instead bounded, frontier-level capabilities proliferate through meek models into the hands of the many.

AI评估算力经济模型公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。