聚焦幻觉发生位置,用偏好学习减少多模态模型幻觉
Stop learning it all to mitigate visual hallucination, Focus on the hallucination target
- 通过定位幻觉目标区域,只在相关部分进行偏好学习
- 在多个任务上显著降低幻觉率,且不损害整体性能
- 适合需要高准确性的视觉-语言应用,如医疗或自动驾驶
多模态大语言模型在视觉-语言任务中常产生幻觉,即生成输入图像中不存在的物体信息,严重影响实际应用中的可靠性。为此,我们提出 \mymethod\,一种聚焦幻觉目标区域的偏好学习方法。构建包含幻觉响应、正确响应及目标信息(图像中真实存在的物体及其在回答中对应的片段位置)的数据集,仅对这些特定目标区域应用偏好学习,使模型过滤无关信号,专注于修正幻觉。实验表明,该方法在多个视觉幻觉任务中有效降低幻觉,提升模型事实性与可靠性,同时保持整体性能不变。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) frequently suffer from hallucination issues, generating information about objects that are not present in input images during vision-language tasks. These hallucinations particularly undermine model reliability in practical applications requiring accurate object identification. To address this challenge, we propose \mymethod,\ a preference learning approach that mitigates hallucinations by focusing on targeted areas where they occur. To implement this, we build a dataset containing hallucinated responses, correct responses, and target information (i.e., objects present in the images and the corresponding chunk positions in responses affected by hallucinations). By applying a preference learning method restricted to these specific targets, the model can filter out irrelevant signals and focus on correcting hallucinations. This allows the model to produce more factual responses by concentrating solely on relevant information. Experimental results demonstrate that \mymethod\ effectively reduces hallucinations across multiple vision hallucination tasks, improving the reliability and performance of MLLMs without diminishing overall performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。