arXiv:2512.02231cs.CVcs.AI2025-12被引 16

评测大模型对语音、视觉和语言的联合理解能力,聚焦说话人识别与时空对齐。

See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models

  • 以说话人为核心设计多模态推理任务,强调跨模态对齐。
  • 3212道题测试模型在真实视频中分辨谁在何时说什么的能力。
  • 发现闭源模型表现远超开源模型,差距主要来自音频视觉融合不足。

多模态大语言模型(MLLMs)应能联合解析视觉、音频与语言信息,但现有视频基准测试很少评估人类语音的细粒度推理。许多任务仅靠视觉即可解决,或粗略评估语音内容,难以判断模型是否真正理解说话人、所说内容及发生时间。我们提出 AV-SpeakerBench,一个包含 3,212 道选择题的精选基准,聚焦真实视频中的以说话人为中心的多模态推理。该基准具有三个特点:(1) 以说话人为核心推理单元,而非场景;(2) 问题设计融合视听依赖关系,嵌入语义;(3) 经专家标注,确保时间精度与跨模态一致性。全面评估显示,Gemini 系列持续优于开源系统,其中 Gemini 2.5 Pro 表现最佳。在开源模型中,Qwen3-Omni-30B 接近 Gemini 2.0 Flash,但与 Gemini 2.5 Pro 相差甚远,主因是音频视觉融合能力弱于视觉感知。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) are expected to jointly interpret vision, audio, and language, yet existing video benchmarks rarely assess fine-grained reasoning about human speech. Many tasks remain visually solvable or only coarsely evaluate speech, offering limited insight into whether models can align who speaks, what is said, and when it occurs. We introduce AV-SpeakerBench, a curated benchmark of 3,212 multiple-choice questions focused on speaker-centric audiovisual reasoning in real-world videos. It features: (1) a speaker-centered formulation that treats speakers-not scenes-as the core reasoning unit; (2) fusion-grounded question design embedding audiovisual dependencies into question semantics; and (3) expert-curated annotations ensuring temporal precision and cross-modal validity. Comprehensive evaluations show that the Gemini family consistently outperforms open-source systems, with Gemini 2.5 Pro achieving the best results. Among open models, Qwen3-Omni-30B approaches Gemini 2.0 Flash but remains far behind Gemini 2.5 Pro, primarily due to weaker audiovisual fusion rather than visual perception. We believe AV-SpeakerBench establishes a rigorous foundation for advancing fine-grained audiovisual reasoning in future multimodal systems.

多模态语音理解推理评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。