arXiv:2604.02605cs.AIcs.SD2026-04被引 7

AVLLM虽能听懂声音,但遇到视觉冲突时仍偏信视觉。

Do Audio-Visual Large Language Models Really See and Hear?

  • 通过层间分析揭示音视频特征融合机制
  • 音频信息在深层被视觉压制,导致输出偏差
  • 训练数据不足致音频对齐弱,适合研究多模态偏差者

视听大语言模型(AVLLMs)正成为统一的多模态感知接口。本文首次开展AVLLM的机械可解释性研究,分析音视频特征如何在不同网络层中演变与融合,最终生成文本输出。结果表明,尽管中间层编码了丰富的音频语义,但当音频与视觉冲突时,这些能力在最终文本生成中基本失效。探针分析显示,有用音频信息仍存在于潜在空间,但深层融合层过度偏好视觉表示,抑制了音频线索。进一步追踪发现,该不平衡源于训练:AVLLM的音频行为强烈匹配其视觉-语言基模型,表明对音频监督的额外对齐有限。研究揭示了AVLLMs中的根本模态偏见,并为多模态大模型如何整合音视频提供了新的机制洞察。

原文摘要 · Abstract (English)

Audio-Visual Large Language Models (AVLLMs) are emerging as unified interfaces to multimodal perception. We present the first mechanistic interpretability study of AVLLMs, analyzing how audio and visual features evolve and fuse through different layers of an AVLLM to produce the final text outputs. We find that although AVLLMs encode rich audio semantics at intermediate layers, these capabilities largely fail to surface in the final text generation when audio conflicts with vision. Probing analyses show that useful latent audio information is present, but deeper fusion layers disproportionately privilege visual representations that tend to suppress audio cues. We further trace this imbalance to training: the AVLLM's audio behavior strongly matches its vision-language base model, indicating limited additional alignment to audio supervision. Our findings reveal a fundamental modality bias in AVLLMs and provide new mechanistic insights into how multimodal LLMs integrate audio and vision.

多模态可解释性模型偏见

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。