用模型间分歧补足自一致性,更好识别大模型的自信错误。
Complementing Self-Consistency with Cross-Model Disagreement for Uncertainty Quantification

- 引入跨模型语义分歧度量,作为黑盒场景下的认知不确定性指标。
- 在10个长文本任务中,总不确定性使排名校准更准,能有效识别自信错误。
- 适合需要可靠置信度评估的高风险应用,如医疗或金融决策。
大语言模型常产生自信但错误的回答,不确定性量化是提升鲁棒性的关键。现有方法依赖自一致性估计随机不确定性(AU),但当模型过度自信并重复输出相同错误答案时,该指标失效。我们分析此情形发现:错误答案下跨模型语义分歧更高,而此时AU偏低。受此启发,我们提出一种在黑盒访问下可用的认知不确定性(EU)项——仅需一个规模匹配的小型模型集成生成文本,通过模型间与模型内序列语义相似性的差距计算。总不确定性(TU)定义为AU与EU之和。在五个7-9B参数量的指令微调模型及十个长文本任务上的全面实验表明,相比仅用AU,TU提升了排名校准性能与选择性拒答效果;EU能可靠标记出AU低但模型高度自信的失败案例。我们进一步通过一致性和互补性诊断刻画了EU最有效的场景。
原文摘要 · Abstract (English)
Large language models (LLMs) often produce confident yet incorrect responses, and uncertainty quantification is one potential solution to more robust usage. Recent works routinely rely on self-consistency to estimate aleatoric uncertainty (AU), yet this proxy collapses when models are overconfident and produce the same incorrect answer across samples. We analyze this regime and show that cross-model semantic disagreement is higher on incorrect answers precisely when AU is low. Motivated by this, we introduce an epistemic uncertainty (EU) term that operates in the black-box access setting: EU uses only generated text from a small, scale-matched ensemble and is computed as the gap between inter-model and intra-model sequence-semantic similarity. We then define total uncertainty (TU) as the sum of AU and EU. In a comprehensive study across five 7-9B instruction-tuned models and ten long-form tasks, TU improves ranking calibration and selective abstention relative to AU, and EU reliably flags confident failures where AU is low. We further characterize when EU is most useful via agreement and complementarity diagnostics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。