arXiv:2506.16724cs.CLcs.AI2025-06EMNLP被引 1

模型置信度越低,偏差对不确定性估计的干扰越大。

The Role of Model Confidence on Bias Effects in Measured Uncertainties for Vision-Language Models

  • 分析不同置信度下偏差对视觉语言模型不确定性估计的影响
  • 低置信度时偏差导致认知不确定性被严重低估
  • 适合关注大模型可靠性与偏差控制的研究者

随着大语言模型在开放任务中的广泛应用,准确评估反映模型知识缺乏的认知不确定性变得至关重要。然而,由于存在多个合理答案带来的偶然不确定性,量化认知不确定性仍具挑战。本文在视觉问答任务中研究偏差对不确定性估计的影响,发现降低提示引入的偏差可提升GPT-4o的不确定性量化效果。基于前人发现:当模型置信度低时倾向于复制输入信息,我们进一步分析了不同无偏置置信度水平下,提示偏差对认知与偶然不确定性的影响。结果表明,所有考虑的偏差在无偏置置信度较低时对两类不确定性影响更显著;且低置信度时偏差会导致认知不确定性被系统性低估,造成过度自信,但对偶然不确定性估计的方向性影响不显著。该发现深化了对偏差缓解与不确定性量化关系的理解,或可指导更先进方法的发展。

原文摘要 · Abstract (English)

With the growing adoption of Large Language Models (LLMs) for open-ended tasks, accurately assessing epistemic uncertainty, which reflects a model's lack of knowledge, has become crucial to ensuring reliable outcomes. However, quantifying epistemic uncertainty in such tasks is challenging due to the presence of aleatoric uncertainty, which arises from multiple valid answers. While bias can introduce noise into epistemic uncertainty estimation, it may also reduce noise from aleatoric uncertainty. To investigate this trade-off, we conduct experiments on Visual Question Answering (VQA) tasks and find that mitigating prompt-introduced bias improves uncertainty quantification in GPT-4o. Building on prior work showing that LLMs tend to copy input information when model confidence is low, we further analyze how these prompt biases affect measured epistemic and aleatoric uncertainty across varying bias-free confidence levels with GPT-4o and Qwen2-VL. We find that all considered biases have greater effects in both uncertainties when bias-free model confidence is lower. Moreover, lower bias-free model confidence is associated with greater bias-induced underestimation of epistemic uncertainty, resulting in overconfident estimates, whereas it has no significant effect on the direction of bias effect in aleatoric uncertainty estimation. These distinct effects deepen our understanding of bias mitigation for uncertainty quantification and potentially inform the development of more advanced techniques.

不确定性估计大模型偏差视觉问答置信度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。