arXiv:2603.07048cs.CVcs.AI2026-03被引 1

通过跨图注意力校准与偏好学习,减少多图生成中的幻觉问题。

Looking Back and Forth: Cross-Image Attention Calibration and Attentive Preference Learning for Multi-Image Hallucination Mitigation

  • 设计可选图像标记交互注意力,实现细粒度跨图对齐。
  • 对比全交互与互不可见的推理结果,提升视觉证据依赖性。
  • 在多图任务中显著降幻觉,且单图任务性能不下降。

尽管大视觉语言模型(LVLMs)展现出强大能力,但在多图任务中仍易产生幻觉。我们将其归因于现有注意力机制的局限性和跨图建模不足。为此,提出结构化幻觉缓解框架CAPL,包含跨图注意力校准与偏好学习。CAPL在架构层面增强跨图交互,并在训练中强化对真实跨图证据的依赖,从而提升模型对跨图关联的感知与建模能力。具体而言:(i) 引入可选图像标记交互注意力机制,实现细粒度跨图实体对齐与信息流动;(ii) 设计基于跨图建模的偏好优化策略,对比全交互与互不可见时的推理结果,促使模型基于真实视觉证据进行预测,减少受文本先验驱动的错误推断。实验表明,CAPL在多种模型架构上均持续提升性能,多图幻觉与通用基准均有稳定增益。单图视觉任务性能保持稳定或略有提升,体现良好泛化能力。

原文摘要 · Abstract (English)

Although large vision-language models (LVLMs) have demonstrated remarkable capabilities, they are prone to hallucinations in multi-image tasks. We attribute this issue to limitations in existing attention mechanisms and insufficient cross-image modeling. Inspired by this, we propose a structured hallucination mitigation framework involving Cross-Image Attention calibration and Preference Learning (CAPL). CAPL explicitly enhances inter-image interactions at the architectural level while reinforcing reliance on genuine cross-image evidence during training, thereby improving the model's perception and modeling of cross-image associations. Specifically, we (i) introduce a selectable image token interaction attention mechanism to establish fine-grained cross-image entity alignment and information flow; (ii) design a cross-image modeling-based preference optimization strategy that contrasts reasoning outcomes under full inter-image interaction and those obtained when images are mutually invisible, encouraging the model to ground its predictions in authentic visual evidence and mitigating erroneous inferences driven by textual priors. Experimental results demonstrate that CAPL consistently improves performance across multiple model architectures, achieving stable gains on both multi-image hallucination and general benchmarks. Notably, performance on single-image visual tasks remains stable or slightly improves, indicating strong generalization capability.

幻觉抑制跨图建模注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。