arXiv:2604.12951cs.LG2026-04被引 1

AI模型越准,越难验证其可靠性,这是根本限制。

The Verification Tax: Fundamental Limits of AI Auditing in the Rare-Error Regime

  • 证明校准误差估计存在理论极限,无法突破
  • 模型越精确,验证难度指数级上升,23%对比无法区分
  • 自评估无效,需主动查询才能可靠检测偏差

深度学习中最受引用的校准结果——在CIFAR-100上温度后调整的ECE为0.012(Guo等,2017)——低于统计噪声下限。我们证明这并非实验缺陷,而是一种规律:以模型错误率ε为条件,校准误差估计的极小极大率为Θ((Lε/m)^{1/3}),任何估计器都无法超越此界。这一“验证税”表明,随着模型性能提升,验证其校准性变得愈发根本困难——计算方向相反却指数级加剧。我们建立四项颠覆常规评估实践的结果:(1) 无标签自评估对校准毫无信息量,上限恒定且与算力无关;(2) 在mε≈1处出现尖锐相变,低于该值时偏差不可检测;(3) 主动查询可消除利普希茨常数,使估计退化为检测;(4) 验证成本随流水线深度以速率L^K指数增长。我们在五个基准(MMLU、TruthfulQA、ARC-Challenge、HellaSwag、WinoGrande;约27,000项)上验证,涵盖6个大语言模型(5个系列,参数规模8B–405B),共27个模型-基准组合,使用基于logprob的置信度,95%自助置信区间与置换检验。自评估不显著在80%组合中成立。在前沿模型中,23%的成对比较无法区别于噪声,意味着可信的校准声明必须报告验证下限,并在性能接近基准分辨能力时优先采用主动查询。

原文摘要 · Abstract (English)

The most cited calibration result in deep learning -- post-temperature-scaling ECE of 0.012 on CIFAR-100 (Guo et al., 2017) -- is below the statistical noise floor. We prove this is not a failure of the experiment but a law: the minimax rate for estimating calibration error with model error rate epsilon is Theta((Lepsilon/m)^{1/3}), and no estimator can beat it. This "verification tax" implies that as AI models improve, verifying their calibration becomes fundamentally harder -- with the same exponent in opposite directions. We establish four results that contradict standard evaluation practice: (1) self-evaluation without labels provides exactly zero information about calibration, bounded by a constant independent of compute; (2) a sharp phase transition at mepsilon approx 1 below which miscalibration is undetectable; (3) active querying eliminates the Lipschitz constant, collapsing estimation to detection; (4) verification cost grows exponentially with pipeline depth at rate L^K. We validate across five benchmarks (MMLU, TruthfulQA, ARC-Challenge, HellaSwag, WinoGrande; ~27,000 items) with 6 LLMs from 5 families (8B-405B parameters, 27 benchmark-model pairs with logprob-based confidence), 95% bootstrap CIs, and permutation tests. Self-evaluation non-significance holds in 80% of pairs. Across frontier models, 23% of pairwise comparisons are indistinguishable from noise, implying that credible calibration claims must report verification floors and prioritize active querying once gains approach benchmark resolution.

AI校准验证极限大模型评估统计噪声

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。