arXiv:2509.17901cs.CVcs.MM2025-09中稿 · Interspeech 2026

视频大模型常忽视音频,新基准发现听觉对理解至关重要

Do Modern Video-LLMs Need to Listen? A Benchmark Audit and Scalable Remedy

  • 在10个基准上测试,仅靠视觉就能答对76%的视听问答
  • 加入语音编码器后,在需语音理解的任务中性能显著提升
  • 适合关注多模态融合与真实场景理解的研究者

多年来发展成熟的语音与音频编码器常被排除在视频理解流程之外,并非因其失效,而是因现有基准从未要求模型‘听’。我们审计了10个视频基准,发现其中多数问题仅凭视觉线索即可解决:单帧探测在视听问答(AVQA)中可正确回答约76%的问题,表明当前评估体系未能有效衡量视听推理能力。基于LLaVA-OneVision,我们引入语音/音频编码器,并在25倍的令牌压缩率下(25 Hz降至1 Hz)对比五种压缩架构。在10个基准上,无论是否过滤,涉及语音理解或跨模态对齐的任务均明显受益于音频输入;而以视觉为中心的任务则基本不受影响。结果表明,语音编码器在视频理解中的作用远超现有基准所反映的程度。相关代码与数据将开源至https://github.com/naver-ai/unimambamia-av。

原文摘要 · Abstract (English)

Speech and audio encoders developed over years of community effort are routinely excluded from video understanding pipelines, not because they fail, but because benchmarks never required listening. We audit 10 video benchmarks and find items largely solvable from visual cues alone: a single-frame probe answers about 76% of AVQA without audio, suggesting poor measurement of audio-visual reasoning. Building on LLaVA-OneVision, we attach a speech/audio encoder and compare five compressor architectures under 25-fold token reduction (25 Hz to 1 Hz). Across 10 benchmarks, with and without filtering, audio yields clear gains on tasks requiring speech comprehension or cross-modal grounding, while vision-centric suites remain largely unaffected. Our results show that speech encoders play a larger role in video understanding than current benchmarks suggest. We will open-source our work at https://github.com/naver-ai/unimambamia-av.

视频理解多模态语音编码基准评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。