arXiv:2506.11073cs.CLcs.AI2025-06ACL被引 16

通过干预跨语言注意力,有效减少多语言视觉模型的幻觉问题。

CLAIM: Mitigating Multilingual Object Hallucination in Large Vision-Language Models with Cross-Lingual Attention Intervention

  • 利用跨语言注意力模式差异,设计近零训练的修正方法
  • 在多个语言上提升准确率,最高改善达30%(西班牙语)
  • 适用于希望低成本优化多语言视觉模型的开发者

大型视觉语言模型虽具强大多模态能力,但在非英语查询时更易产生与图像不符的物体幻觉。现有方法多依赖预训练或微调,成本高昂。本文受跨模态注意力模式语言差异启发,提出一种近零训练的修正方法CLAIM,通过识别语言特异性注意力头、估计英到目标语言的偏移向量,并在推理阶段干预注意力输出,实现跨语言视觉感知对齐。大量实验表明,CLAIM在POPE基准平均提升13.56%(西班牙语最高达30%),在MME幻觉子集上提升21.75%。分析显示,中层注意力在多语言场景下差异最显著,是关键干预位置。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have demonstrated impressive multimodal abilities but remain prone to multilingual object hallucination, with a higher likelihood of generating responses inconsistent with the visual input when utilizing queries in non-English languages compared to English. Most existing approaches to address these rely on pretraining or fine-tuning, which are resource-intensive. In this paper, inspired by observing the disparities in cross-modal attention patterns across languages, we propose Cross-Lingual Attention Intervention for Mitigating multilingual object hallucination (CLAIM) in LVLMs, a novel near training-free method by aligning attention patterns. CLAIM first identifies language-specific cross-modal attention heads, then estimates language shift vectors from English to the target language, and finally intervenes in the attention outputs during inference to facilitate cross-lingual visual perception capability alignment. Extensive experiments demonstrate that CLAIM achieves an average improvement of 13.56% (up to 30% in Spanish) on the POPE and 21.75% on the hallucination subsets of the MME benchmark across various languages. Further analysis reveals that multilingual attention divergence is most prominent in intermediate layers, highlighting their critical role in multilingual scenarios.

视觉语言模型多语言注意力干预幻觉抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。