通过生成幻觉诱导图像,精准定位并抑制视觉语言模型的幻觉问题。
HII-DPO: Eliminate Hallucination via Accurate Hallucination-Inducing Counterfactual Images
- 设计新流程合成幻觉诱导图像(HIIs),揭示场景依赖的幻觉模式。
- 在标准基准上实现38%的幻觉减少,优于现有最优方法。
- 适合关注模型可靠性与多模态对齐的研究者使用。
大型视觉语言模型在多种多模态任务中表现卓越,但易受固有语言偏见引发的幻觉影响。现有缓解方法常忽略由语言偏见驱动的幻觉模式。本文提出新流程,精准合成幻觉诱导图像(HIIs)。利用这些图像,我们发现模型在无视觉证据时仍倾向于提及与场景高度相关的典型物体,呈现一致的场景依赖幻觉模式。为量化模型对此类幻觉的敏感度,我们构建了掩码物体幻觉(MOH)基准,严格评估现有先进对齐框架。最后,利用HIIs构建高质量偏好数据集,实现细粒度对齐。实验表明,该方法有效降低幻觉,同时保持模型通用能力,在标准幻觉基准上相比当前最优方法提升最高达38%。
原文摘要 · Abstract (English)
Large Vision-Language Models (VLMs) have achieved remarkable success across diverse multimodal tasks but remain vulnerable to hallucinations rooted in inherent language bias. Despite recent progress, existing hallucination mitigation methods often overlook the underlying hallucination patterns driven by language bias. In this work, we design a novel pipeline to accurately synthesize Hallucination-Inducing Images (HIIs). Using synthesized HIIs, we reveal a consistent scene-conditioned hallucination pattern: models tend to mention objects that are highly typical of the scene even when visual evidence is removed. To quantify the susceptibility of VLMs to this hallucination pattern, we establish the Masked-Object-Hallucination (MOH) benchmark to rigorously evaluate existing state-of-the-art alignment frameworks. Finally, we leverage HIIs to construct high-quality preference datasets for fine-grained alignment. Experimental results demonstrate that our approach effectively mitigates hallucinations while preserving general model capabilities. Specifically, our method achieves up to a 38% improvement over the current state-of-the-art on standard hallucination benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。