arXiv:2605.08145cs.CVcs.AI2026-05中稿 · ICML

通过增强模态间冗余信息,提升视觉语言模型的鲁棒性与一致性。

Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language Models

论文配图:Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language Models
图 1 · 摘自论文原文
  • 设计自描述流程,将独特信息转化为共享冗余信息。
  • 使视觉错误降低38.3%,一致性提升16.8%。
  • 适合关注模型鲁棒性与多模态融合的研究者。

当前视觉语言模型在面对模糊或受损模态时存在幻觉和鲁棒性问题。我们假设可通过利用模态间的共享信息来弥补受损模态。为此,分析了多模态交互——模态提供的冗余(共享)、独特(独占)和协同(涌现)的任务相关信息,以确定其对模型可靠性的影响。具体而言,放大冗余交互可增加可利用的共享信息,从而缓解这些问题;然而现代指令数据集常消除冗余以强调视觉定位。我们通过自描述工作流引入多模态交互门机制,将独特交互转化为冗余交互。实验表明,增加冗余可使视觉诱发误差减少38.3%,一致性提升16.8%。

原文摘要 · Abstract (English)

Current vision language models face hallucination and robustness issues against ambiguous or corrupted modalities. We hypothesize that these issues can be addressed by exploiting the shared information between modalities to compensate for the impaired one. To this end, we analyze multimodal interactions -- redundant (shared), unique (exclusive), and synergistic (emergent) task-relevant information provided by the modalities -- to determine their impacts on model reliability. Specifically, amplifying redundant interactions would increase this exploitable shared information to resolve these issues; yet, modern instruction datasets often eliminate redundancies to prioritize visual grounding. We bridge this gap through a self-captioning workflow featuring a \textsc{Multimodal Interaction Gate}: a mechanism to convert unique interactions into redundant interactions. Our findings suggest that increasing redundancy can reduce visual induced errors by 38.3\% and improve consistency by 16.8\%.

多模态视觉语言鲁棒性冗余增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。