让视觉语言模型在不确定时主动获取更多图像证据,减少盲目拒绝
Look Again Before You Abstain:Budgeted Conformal Evidence Acquisition for Reliable Vision-Language Model

- 用三选一决策替代二选一:回答、拒绝或申请额外图像证据
- 在保持幻觉率低于5%前提下,将拒绝率从80%以上降至40%以下
- 适合需要高可靠性的医疗、自动驾驶等视觉判断场景
大型视觉语言模型会生成图像不支持的视觉细节。一种合理对策是选择性预测并提供无分布保证:验证每个断言,若不成立则拒绝,从而确保断言中的幻觉率可被严格控制。然而我们发现,这种保证代价巨大:为使平衡物体存在性基准上的幻觉率低于5%,现有最先进的共形过滤器需拒绝超过80%的断言。我们认为,当额外视觉证据成本低廉时,拒绝是浪费,因此提出预算内共形证据获取(BCEA),将二元答案/拒绝决策改为三元选择:回答、拒绝或在有限预算内通过重看图像(如缩放、裁剪或特定干预)获取更多证据。
原文摘要 · Abstract (English)
Large vision-language models (LVLMs) hallucinate: they assert visual details that the image does not support. A principled remedy is selective prediction with a distribution-free guarantee-verify each claim and abstain when the claim is not grounded, so that the hallucination rate among asserted claims is provably bounded. We show, however, that this guarantee is bought at a brutal price: to keep the hallucination rate below $5\%$ on a balanced object-existence benchmark, a state-of-the-art conformal filter must abstain on more than $80\%$ of claims. We argue that abstention is wasteful when more visual evidence is cheaply available, and introduce Budgeted Conformal Evidence Acquisition (BCEA), which replaces the binary answer/abstain decision with a three-way choice: answer, abstain, or acquire additional visual evidence by re-examining the image (zooming, cropping, or applying a claim-specific intervention) under a bounded
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。