arXiv:2506.04039cs.CVcs.AI2025-06EMNLP被引 8

通过实体中心对齐,大幅减少视觉语言模型的幻觉问题。

Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference Optimization

  • 以实体为中心构建多模态偏好优化,增强图像与文本对齐
  • 在多个基准上降低幻觉率,最高达85.9%
  • 适用于需要高可信度生成的视觉语言应用

大型视觉语言模型(LVLM)在多项任务中表现出色,但其可信度常受幻觉困扰,这源于模态错位及底层大语言模型(LLM)固有的幻觉。现有偏好对齐方法侧重于匹配人类偏好,忽视图像-文本模态对齐,导致模型过度依赖LLM并产生幻觉。本文提出实体中心多模态偏好优化(EMPO),相比现有方法显著提升模态对齐能力。为解决高质量多模态偏好数据稀缺问题,我们利用开源指令数据集,自动构建涵盖图像、指令和回复三个维度的高质量偏好数据。在两个人类偏好数据集和五个多模态幻觉基准上的实验表明,EMPO效果显著,例如在Object-HalBench上将幻觉率降低85.9%,在MM-HalBench上降低49.8%。

原文摘要 · Abstract (English)

Large Visual Language Models (LVLMs) have demonstrated impressive capabilities across multiple tasks. However, their trustworthiness is often challenged by hallucinations, which can be attributed to the modality misalignment and the inherent hallucinations of their underlying Large Language Models (LLMs) backbone. Existing preference alignment methods focus on aligning model responses with human preferences while neglecting image-text modality alignment, resulting in over-reliance on LLMs and hallucinations. In this paper, we propose Entity-centric Multimodal Preference Optimization (EMPO), which achieves enhanced modality alignment compared to existing human preference alignment methods. Besides, to overcome the scarcity of high-quality multimodal preference data, we utilize open-source instruction datasets to automatically construct high-quality preference data across three aspects: image, instruction, and response. Experiments on two human preference datasets and five multimodal hallucination benchmarks demonstrate the effectiveness of EMPO, e.g., reducing hallucination rates by 85.9\% on Object-HalBench and 49.8\% on MM-HalBench.

视觉语言模型幻觉抑制多模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。