arXiv:2606.30814cs.CL2026-06被引 1

提出新评估框架ACE,让大模型校准能力比较更公平

When Calibration Rankings Reverse: Accuracy-Controlled Evaluation for Fair Comparison of LLMs

  • 用三种对齐视角控制准确率,避免误差干扰
  • 发现多数模型校准优势在控准后大幅减弱
  • 适合想公平对比大模型可靠性的研究者

校准评估模型置信度与实际准确率的一致性。现有研究多用全局校准指标(如期望校准误差、Brier Score)比较不同大语言模型的校准性能。我们从理论和实证两方面证明,此类比较受模型准确率差异的干扰。为此,提出ACE——一种包含实例对齐、分布对齐和候选对齐三个互补视角的准确率控制评估框架。在多个基准测试、模型族和置信度生成方法下,利用ACE研究小模型与大模型、思考型与非思考型模型两类重要对比维度。结果表明,许多先前在原始全局指标下报告的校准优势,在准确率控制后显著削弱;且排名反转现象频繁出现:原本占优的模型在控准后不再领先。结果表明,原始全局校准指标不适用于跨模型比较,公平的校准比较必须考虑准确率影响。

原文摘要 · Abstract (English)

Calibration evaluates whether a model confidence aligns with its empirical accuracy. Existing studies often compare the calibration of different large language models using global calibration metrics such as Expected Calibration Error and Brier Score. We begin by showing, both theoretically and empirically, that such comparisons are confounded by differences in model accuracy. For fairer cross-model comparison, we then propose ACE, an accuracy-controlled evaluation framework with three complementary views: Instance-Aligned, Distribution-Aligned, and Candidate-Aligned calibration. Across multiple benchmarks, model families, and confidence elicitation methods, we use ACE to study two practically important comparison axes, small versus large models and thinking versus non-thinking models. We find that many previously reported calibration advantages under raw global metrics weaken substantially after accuracy control. We also find that ranking reversal is frequent: models favored by raw metrics often cease to be favored once accuracy is controlled. Our results show that raw global calibration metrics are not robust for cross-model comparison, and that fair calibration comparison requires accuracy-aware evaluation.

大模型评估校准公平比较

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。