分析物理AI评测的冗余问题,发现12个基准测试信息重叠严重。
A Statistical Audit of Physical AI Benchmark Redundancy

- 构建51模型×12基准矩阵,量化评测间信息共享程度
- 发现冗余导致22个模型排名变动超3位,影响评估可靠性
- 提出贪心筛选法,4个基准保留78.5%原始信息量
物理AI模型在不同报告中采用各异的评测套件,导致模型-评测矩阵稀疏且评测间关系不明。本文从51个评测和152个模型中,依据报告密度选取51个模型与12个物理AI评测,整合模型卡片与评测论文得分,并基于官方协议自行评估。通过测量评测间信息共享度,提供冗余的量化证据:将两组可互换评测合并后,51个模型中有22个在等权重平均排名中变动3位及以上。随后采用贪心策略,在兼顾得分离散度与未解释方差的前提下筛选基准,得到仅含4个评测的子集,其保留了全部12个评测78.5%的效用。该方法仅需评测级得分且有足够重叠,不局限于物理AI领域。
原文摘要 · Abstract (English)
Physical AI models are evaluated on suites of benchmarks that differ across model reports, leaving the model-by-benchmark matrix sparse and the relationship between benchmarks unmeasured. We construct a matrix of 51 models on 12 physical AI benchmarks, selected from a registry of 51 benchmarks and 152 models by reporting density, combining scores from model cards and benchmark papers with our own evaluation runs under each benchmark's official protocol. We measure how much information the benchmarks share and show quantitative evidence of Redundancy. Redundancy affects reported rankings: collapsing the two substitute pairs into single columns moves 22 of 51 models by three or more places under an equally weighted average. We then select benchmarks greedily under a utility combining score dispersion with variance not explained by the already-selected set, and obtain a four-benchmark subset retaining 78.5\% of the utility of all 12, on which we fit a Bradley--Terry ranking. The procedure requires only benchmark-level scores with sufficient overlap and is not specific to physical AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。