arXiv:2503.06486cs.CVcs.AI2025-03ICLR被引 48

通过扰动文本训练,让多模态模型更依赖视觉信息,减少幻觉。

PerturboLLaVA: Reducing Multimodal Hallucinations with Perturbative Visual Training

  • 用对抗性文本扰动训练,降低模型对语言先验的依赖。
  • 提出新评估指标HalFscore,细粒度衡量图像描述准确性和完整性。
  • 无需额外计算开销,显著减少幻觉,适合图像描述任务使用。

本文针对多模态大模型在密集图像描述任务中存在幻觉的问题,发现当前缺乏能在概念层面精细评估描述质量的指标。为此,提出基于语言图的全新指标HalFscore,可细粒度评估密集描述的准确性和完整性。同时,识别出幻觉根源为模型过度依赖语言先验。为此,提出PerturboLLaVA方法,在训练中引入对抗性扰动文本,增强模型对视觉输入的关注,有效减少幻觉,生成更符合图像的事实性描述,且不增加额外计算开销。该方法在处理多模态幻觉方面优于现有方法,并在通用多模态基准上实现性能提升。

原文摘要 · Abstract (English)

This paper aims to address the challenge of hallucinations in Multimodal Large Language Models (MLLMs) particularly for dense image captioning tasks. To tackle the challenge, we identify the current lack of a metric that finely measures the caption quality in concept level. We hereby introduce HalFscore, a novel metric built upon the language graph and is designed to evaluate both the accuracy and completeness of dense captions at a granular level. Additionally, we identify the root cause of hallucination as the model's over-reliance on its language prior. To address this, we propose PerturboLLaVA, which reduces the model's reliance on the language prior by incorporating adversarially perturbed text during training. This method enhances the model's focus on visual inputs, effectively reducing hallucinations and producing accurate, image-grounded descriptions without incurring additional computational overhead. PerturboLLaVA significantly improves the fidelity of generated captions, outperforming existing approaches in handling multimodal hallucinations and achieving improved performance across general multimodal benchmarks.

多模态幻觉抑制图像描述训练方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。