揭示多模态输入冲突导致模型幻觉,提出三类缓解方法。
Robust Multimodal Large Language Models Against Modality Conflict

- 从模态冲突角度分析幻觉成因,构建新数据集MMMC。
- 强化学习法在冲突场景下抑制幻觉效果最佳,微调法稳定可靠。
- 适合关注多模态模型鲁棒性的研究者和开发者参考。
尽管多模态大语言模型(MLLMs)在视觉-语言任务中表现优异,但在真实场景中仍易产生幻觉。本文从模态冲突视角研究这一现象,不同于以往关注模型输出与输入的矛盾,我们聚焦于不同模态输入间的内在冲突,这类冲突使模型陷入困境并直接引发幻觉。我们正式定义了模态冲突,并构建了名为Multimodal Modality Conflict(MMMC)的数据集以模拟该现象。针对此问题,提出了基于提示工程、监督微调和强化学习三种方法。在MMMC数据集上开展大量实验,分析各方法的优劣。结果表明,强化学习方法在缓解模态冲突引起的幻觉方面表现最佳,而监督微调方法展现出良好且稳定的性能。本工作揭示了导致幻觉的未被注意的模态冲突,为提升MLLMs的鲁棒性提供了新洞见。
原文摘要 · Abstract (English)
Despite the impressive capabilities of multimodal large language models (MLLMs) in vision-language tasks, they are prone to hallucinations in real-world scenarios. This paper investigates the hallucination phenomenon in MLLMs from the perspective of modality conflict. Unlike existing works focusing on the conflicts between model responses and inputs, we study the inherent conflicts in inputs from different modalities that place MLLMs in a dilemma and directly lead to hallucinations. We formally define the modality conflict and construct a dataset named Multimodal Modality Conflict (MMMC) to simulate this phenomenon in vision-language tasks. Three methods based on prompt engineering, supervised fine-tuning, and reinforcement learning are proposed to alleviate the hallucination caused by modality conflict. Extensive experiments are conducted on the MMMC dataset to analyze the merits and demerits of these methods. Our results show that the reinforcement learning method achieves the best performance in mitigating the hallucination under modality conflict, while the supervised fine-tuning method shows promising and stable performance. Our work sheds light on the unnoticed modality conflict that leads to hallucinations and provides more insights into the robustness of MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。