arXiv:2605.09906cs.AIcs.SD2026-05

提出分步推理框架,减少视听大模型的模态干扰问题。

Separate First, Fuse Later: Mitigating Cross-Modal Interference in Audio-Visual LLMs Reasoning with Modality-Specific Chain-of-Thought

论文配图:Separate First, Fuse Later: Mitigating Cross-Modal Interference in Audio-Visual LLMs Reasoning with Modality-Specific Chain-of-Thought
图 1 · 摘自论文原文
  • 先分后融:分别生成视听独立推理链,最后融合证据
  • 在通用评测上提升5.16%,在幻觉检测上提升11.17%
  • 适合需要高可靠多模态推理的应用场景

音频与视觉为视听问答提供互补信息,但当前视听大语言模型可能因跨模态干扰而产生误判:一模态信息误导另一模态理解,导致幻觉。我们将其归因于中间推理阶段的无控跨模态交互。为此,提出分离先、融合后(SFFL)框架,以降低跨模态干扰。SFFL强制执行模态特异性链式推理,生成独立的音频与视觉推理轨迹,并在最终整合证据作答。通过不同模态输入设置构建模态偏好标签,作为强化学习中的辅助奖励,鼓励实例依赖的模态线索偏好。进一步引入模态特异性推理机制,在分离推理阶段保持模态隔离,而在证据融合阶段实现全量跨模态信息访问。实验显示,在准确率与鲁棒性上均有持续提升,通用AVQA基准平均相对增益达5.16%,跨模态幻觉基准提升11.17%。

原文摘要 · Abstract (English)

Audio and vision provide complementary evidence for audio-visual question answering, yet current audio-visual large language models may suffer from cross-modal interference: information from one modality misguides the interpretation of another, thereby inducing hallucinations. We attribute this issue to uncontrolled cross-modal interactions during intermediate reasoning. To mitigate this, we propose Separate First, Fuse Later (SFFL), an audio-visual reasoning framework designed to reduce cross-modal interference. SFFL enforces modality-specific chain-of-thought reasoning, producing separate audio and visual reasoning traces and integrating evidence for answering. We construct modality-preference labels via a data pipeline under different modality input settings. We use these labels as an auxiliary reward in reinforcement learning to encourage a instance-dependent preference for modality cues when answering. We further introduce a modality-specific reasoning mechanism that preserves modality isolation during the separated reasoning stage while enabling full access to cross-modal information at the evidence fusion stage. Experiments demonstrate consistent improvements in both accuracy and robustness, yielding an average relative gain of 5.16\% on general AVQA benchmarks and 11.17\% on a cross-modal hallucination benchmark.

多模态链式推理幻觉抑制视听模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。