让视觉语言模型自评更准,靠的是看答案多依赖图像证据
VAUQ: Vision-Aware Uncertainty Quantification for LVLM Self-Evaluation
- 用图像信息得分衡量答案对视觉内容的依赖程度
- 在多个数据集上自评准确率显著优于现有方法
- 无需训练,适合提升视觉推理模型可靠性
大型视觉语言模型(LVLM)常出现幻觉,限制其在真实场景中的安全部署。现有大语言模型自评估方法依赖模型对自身输出正确性的判断,虽可提升部署可靠性,但过度依赖语言先验,难以评估视觉条件下的预测。本文提出VAUQ,一种面向LVLM自评估的视觉感知不确定性量化框架,通过显式度量模型输出对视觉证据的依赖强度来改进评估。VAUQ引入图像信息得分(IS),捕捉视觉输入带来的预测不确定性降低;同时采用无监督的核心区域掩码策略,增强显著区域的影响。将预测熵与核心掩码后的IS结合,形成无需训练的评分函数,能可靠反映答案正确性。大量实验表明,VAUQ在多个数据集上持续优于现有自评估方法。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) frequently hallucinate, limiting their safe deployment in real-world applications. Existing LLM self-evaluation methods rely on a model's ability to estimate the correctness of its own outputs, which can improve deployment reliability; however, they depend heavily on language priors and are therefore ill-suited for evaluating vision-conditioned predictions. We propose VAUQ, a vision-aware uncertainty quantification framework for LVLM self-evaluation that explicitly measures how strongly a model's output depends on visual evidence. VAUQ introduces the Image-Information Score (IS), which captures the reduction in predictive uncertainty attributable to visual input, and an unsupervised core-region masking strategy that amplifies the influence of salient regions. Combining predictive entropy with this core-masked IS yields a training-free scoring function that reliably reflects answer correctness. Comprehensive experiments show that VAUQ consistently outperforms existing self-evaluation methods across multiple datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。