arXiv:2605.28123cs.CL2026-05

针对视觉语言模型幻觉,提出按风险选择性启用提示验证。

Risk-aware Selective Prompting for Hallucination Mitigation in Large Vision-Language Models

论文配图:Risk-aware Selective Prompting for Hallucination Mitigation in Large Vision-Language Models
图 1 · 摘自论文原文
  • 根据输入难度动态决定是否使用提示验证,避免错误引入。
  • 在困难输入上提升纠错率,同时不降低简单输入的表现。
  • 无需训练,利用生成前不确定性信号实现智能触发。

基于提示的验证广泛用于缓解大视觉语言模型(LVLM)中的幻觉问题,但其有效性机制尚不清晰。我们系统研究了两种代表性LVLM架构在两个幻觉基准上的验证提示行为,发现该方法具有风险性:纠正效果随输入难度增加,但新引入的错误在各类难度下均持续存在。因此,始终开启提示验证对困难样本有益,却可能损害简单样本表现。分析显示,这种现象与保守输出倾向相关:提示使注意力从视觉标记转向指令标记,并引发中层熵模式变化,表明是指令引导的注意力重分配,而非统一增强的视觉对齐。为此,我们提出无训练的风险感知选择性提示(RSP),利用生成前的不确定性信号决定是否触发验证。RSP在保持基线性能的同时缓解了始终开启提示的退化问题,并揭示不同架构所需的选择信号存在差异。

原文摘要 · Abstract (English)

Prompt-based verification is widely used to mitigate hallucinations in large vision-language models (LVLMs), yet when it helps remains poorly understood. We systematically study verification prompting across two representative LVLM architectures and hallucination benchmarks, and find that it is a risk-bearing intervention: its corrections increase with input difficulty, while newly introduced errors persist across difficulty levels. As a result, always-on prompting helps on hard inputs but offers little benefit -- and can harm -- easier ones. Our analysis further shows that this behavior is associated with a conservative output shift. Verification prompts redistribute attention from visual tokens toward instruction tokens and induce a distinct middle-layer entropy pattern absent in a neutral-prompt control, suggesting instruction-conditioned attention redistribution rather than uniformly improved visual grounding. Motivated by this input-dependent risk, we propose Risk-aware Selective Prompting (RSP), a training-free approach that uses pre-generation uncertainty signals to trigger verification selectively. RSP mitigates the degradation of always-on prompting while preserving baseline performance, and reveals that effective selection signals vary across architectures.

幻觉缓解视觉语言模型提示工程风险感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。