发现语言模型不确定性评估存在长度偏差干扰,影响结果可靠性
Revisiting Uncertainty Quantification Evaluation in Language Models: Spurious Interactions with Response Length Bias Results
- 揭示评估中正确性指标与不确定性方法的共同长度偏差会扭曲结果
- 实证测试显示8种评估方法在4个数据集上均受长度偏差影响
- 建议使用LM-as-a-judge方法以减少偏差,提升评估公平性
语言模型的不确定性量化(UQ)对提升其安全性和可靠性至关重要。现有评估常采用AUROC等指标,衡量负序列概率等UQ方法与ROUGE-L等正确性函数的相关性。我们发现,当UQ方法和正确性函数同时受相同因素(如响应长度)影响时,这种相互偏差会系统性扭曲评估结果。首先,我们形式化证明任何非随机的共同偏差都会导致AUROC排名失真,破坏基准有效性。其次,我们在4个数据集、4个模型、8种UQ方法下,测试了7种广泛使用的正确性函数(包括基于词法、嵌入和LM作为裁判的方法),证实长度偏差在正确性函数与UQ方法间交互,显著影响评估。研究指出,以LM作为裁判的方法长度偏差最小,为更公平的UQ评估提供了可行路径。
原文摘要 · Abstract (English)
Uncertainty Quantification (UQ) in Language Models (LMs) is key to improving their safety and reliability. Evaluations often use metrics like AUROC to assess how well UQ methods (e.g., negative sequence probabilities) correlate with task correctness functions (e.g., ROUGE-L). We show that mutual biases--when both UQ methods and correctness functions are biased by the same factors--systematically distort evaluation. First, we formally prove that any mutual bias non-randomly skews AUROC rankings, compromising benchmark integrity. Second, we confirm this happens empirically by testing 7 widely used correctness functions, from lexical-based and embedding-based metrics to LM-as-a-judge approaches, across 4 datasets x 4 models x 8 UQ methods. Our analysis shows that length biases in correctness functions distort UQ assessments by interacting with length biases in UQ methods. We identify LM-as-a-judge methods as the least length-biased, offering a promising path for a fairer UQ evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。