揭秘视听大模型如何融合声音与图像信息,揭示其内部信息流动规律。
From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs

- 通过追踪多模态输入路径,发现视听信号按任务需求分步传递
- 音频与视觉令牌在信息传递后可被丢弃,不影响甚至提升预测效果
- 结果适用于多种模型规模与任务,为高效推理提供新思路
多模态大语言模型(MLLMs)能听能看,但音频和视觉信号如何在网络中流动并影响最终决策仍不清晰。本文研究视听大语言模型(AVLLMs)内部的信息流,分析在视频类输入与多个交错的视听项目两种场景下的信息路由机制。研究发现,在视听视频输入中,模型沿视觉语言模型(VLMs)与视频大语言模型(VideoLLMs)的顺序路径传递信息,音频与视觉贡献比例与任务对各模态的依赖度一致;而在多交错视听项场景下,信息流向变为并行通道。此外,我们证明一旦音频-视觉令牌的信息被传至语言模型,即可被丢弃,对预测结果影响极小,甚至略有提升,且该现象在多个任务与数据集上具普适性,适用于3B与7B规模的Qwen2.5-Omni及Video-SALMONN2 Plus模型。这些发现首次勾勒出AVLLMs内在视听协同的完整图景,为未来可解释性、模型设计与推理效率的突破奠定基础。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) can listen and see, but how do audio and visual signals actually travel through the network to shape an answer? Despite their growing role in research and real-world applications, the internal pathways through which audio and visual tokens influence the final prediction remain poorly understood. In this study, we examine audio-visual information flow inside Audio-Visual Large Language Models (AVLLMs), tracing how AVLLMs route, utilize, and integrate audio and visual information across two input configurations, audio-visual video and multiple interleaved audio-visual items. We find that for audio-visual video, AVLLMs follow the sequential information flow pathway established for VLMs and VideoLLMs, with audio and visual contribution flowing along this pathway in proportion to the task's reliance on each modality. In settings with multiple interleaved audio-visual items, this routing shifts to different parallel streams. Furthermore, we demonstrate that audio-visual and other token types can be discarded once their information is transferred to LLM, with minimal impact on the model's prediction or even slight improvement, generalizing across multiple tasks and datasets, enabling more efficient inference. These findings hold across multiple models and scales, Qwen2.5-Omni and Video-SALMONN2 Plus at 3B and 7B scales, leading to hypotheses on why these flow structures emerge. Together, these results deliver the first coherent picture of how AVLLMs orchestrate sound and sight inside the network and lay the groundwork for the next wave of interpretability, design, and efficiency advances in audio-visual and broader MLLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。