提出MESA框架,精准抑制幻觉而不破坏生成行为。
Mitigating Entangled Steering in Large Vision-Language Models for Hallucination Reduction
- 分离幻觉相关信号,选择性干预潜在空间
- 显著降低幻觉率,同时保持原有输出分布
- 适合作为插件用于多种视觉语言模型
大型视觉语言模型在跨模态任务中表现卓越,但常出现与图像内容不符的文本生成(幻觉)。现有方法虽能缓解幻觉,却往往改变生成行为,导致输出变短、词元分布偏移,尤其在潜空间控制方法中更明显。我们发现该问题源于纠缠的引导信号:抑制幻觉会无意中干扰模型固有的生成机制。为此,我们提出MESA——一种即插即用的框架,实现对幻觉相关响应的可控、选择性潜空间干预。MESA聚焦于幻觉相关部分,保留原始词元分布,有效减少幻觉且不损害生成行为。在多个生成与判别基准上的大量实验表明,MESA在不同LVLM家族中均持续降低幻觉,优于以往方法。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) have achieved remarkable success across cross-modal tasks but remain hindered by hallucinations, producing textual outputs inconsistent with visual content. Existing methods mitigate hallucinations but often alter generation behavior, resulting in shorter outputs and shifted token distributions, especially in latent space steering approaches. We identify that this issue stems from entangled steering signals, where suppressing hallucinations inadvertently disrupts the model's intrinsic generation behavior. To address this, we propose MESA, an effective plug-and-play framework that performs controlled and selective latent intervention for hallucination mitigation. Specifically, MESA targets hallucination-relevant responses while preserving the model's original token distribution, enabling effective hallucination reduction without compromising generation behavior. Extensive experiments across diverse generative and discriminative benchmarks demonstrate that MESA consistently reduces hallucinations while better preserving generation behavior, outperforming prior methods across multiple LVLM families.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。