针对多模态大模型幻觉问题,提出靶向优化方法提升准确性。
Mitigating Hallucination in Multimodal Large Language Model via Hallucination-targeted Direct Preference Optimization
- 设计三类针对性偏好数据,分别应对视觉不足、长上下文生成和模态冲突
- 在多个评测集上优于主流方法,显著降低幻觉率
- 适合需要高可信度多模态生成的科研与工业应用
多模态大语言模型(MLLMs)存在幻觉问题,制约实际应用。尽管已有研究尝试将直接偏好优化(DPO)应用于MLLMs,但其在缓解幻觉方面效果参差不齐。为此,本文提出幻觉靶向直接偏好优化(HDPO),从幻觉的多样形式与成因出发进行改进。具体构建三类偏好对数据,分别针对:(1) 视觉理解能力不足,(2) 长上下文生成时的偏差,(3) 多模态信息冲突。实验表明,该方法在多个幻觉评估数据集上表现优异,超越多数现有先进方法,验证了其有效性。消融实验与深入分析进一步证实方法优势,并提示通过扩大规模可实现更高性能。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) are known to hallucinate, which limits their practical applications. Recent works have attempted to apply Direct Preference Optimization (DPO) to enhance the performance of MLLMs, but have shown inconsistent improvements in mitigating hallucinations. To address this issue more effectively, we introduce Hallucination-targeted Direct Preference Optimization (HDPO) to reduce hallucinations in MLLMs. Unlike previous approaches, our method tackles hallucinations from their diverse forms and causes. Specifically, we develop three types of preference pair data targeting the following causes of MLLM hallucinations: (1) insufficient visual capabilities, (2) long context generation, and (3) multimodal conflicts. Experimental results demonstrate that our method achieves superior performance across multiple hallucination evaluation datasets, surpassing most state-of-the-art (SOTA) methods and highlighting the potential of our approach. Ablation studies and in-depth analyses further confirm the effectiveness of our method and suggest the potential for further improvements through scaling up.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。