arXiv:2605.11205cs.LGcs.AI2026-05

简单平均法在数据稀疏时失效,项目反应理论能准确恢复真实排名。

The Scaling Law of Evaluation Failure: Why Simple Averaging Collapses Under Data Sparsity and Item Difficulty Gaps, and How Item Response Theory Recovers Ground Truth Across Domains

论文配图:The Scaling Law of Evaluation Failure: Why Simple Averaging Collapses Under Data Sparsity and Item Difficulty Gaps, and How Item Response Theory Recovers Ground Truth Across Domains
图 1 · 摘自论文原文
  • 用项目反应理论(IRT)替代简单平均,提升评估可靠性。
  • 数据稀疏度67%且难度差异大时,简单平均排名相关性仅0.809。
  • 适用于安全关键领域如自动驾驶、药物测试的评估设计。

跨人工智能与安全关键领域的基准评估普遍依赖简单平均。我们证明当两个条件同时存在时,该方法会产生严重误导:(1) 评估矩阵稀疏;(2) 项目难度差异显著。在四个领域(NLP/GLUE、临床药物试验、自动驾驶安全、网络安全)的受控模拟中,简单平均排名与真实排名的斯皮尔曼相关系数ρ从100%覆盖率下的1.000降至67%覆盖率时的0.809(均值超过20次种子实验)。标准双参数逻辑模型(2PL-IRT)在所有条件下保持ρ≥0.996。150组条件扫描显示,排名误差形成以稀疏度S∈[0,0.70]与难度差距D∈[0.5,5.0]交互主导的失败曲面(γ₃=+0.20, t=13.05),而IRT始终维持ρ≥0.993。研究对物理人工智能评估具有重要启示,因其评估矩阵常不完整且难度差距极大。

原文摘要 · Abstract (English)

Benchmark evaluation across AI and safety-critical domains overwhelmingly relies on simple averaging. We demonstrate that this practice produces substantially misleading rankings when two conditions co-occur: (1) the evaluation matrix is sparse and (2) items vary substantially in difficulty. Through controlled simulation experiments across four domains -- NLP (GLUE), clinical drug trials, autonomous vehicle safety, and cybersecurity -- we show that Spearman rank correlation $ρ$ between simple-average rankings and ground-truth rankings degrades from $ρ= 1.000$ at 100% coverage to $ρ= 0.809$ at 67% coverage with high difficulty heterogeneity (mean over 20 seeds). A standard two-parameter logistic (2PL) Item Response Theory (IRT) model maintains $ρ\geq 0.996$ across all conditions. A 150-condition grid sweep over sparsity $S \in [0, 0.70]$ and difficulty gap $D \in [0.5, 5.0]$ confirms that ranking error forms a failure surface with a strong $S \times D$ interaction ($γ_3 = +0.20$, $t = 13.05$), while IRT maintains $ρ\geq 0.993$ throughout. We discuss implications for Physical AI benchmarking, where evaluation matrices are often incomplete and difficulty gaps are extreme.

评估方法项目反应理论数据稀疏排名可靠

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。