构建首个跨模态音频视觉推理评测基准,推动多模态大模型发展
JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
- 基于多模态大模型自动生成需视听协同理解的问答对
- 顶尖多模态模型平均准确率仅65.3%,跨场景推理仍薄弱
- 覆盖五维认知、四类音频、三类场景,适合评估多模态系统
理解视频本质上需要对视觉和听觉信息进行推理。为有效评估能够处理视觉与音频等多模态信息的通用大语言模型(Omni-LLMs),评测基准必须全面涵盖三个关键维度:(1) 多模态依赖性(即仅靠视觉或音频无法回答的问题),(2) 多样化的音频类型(如语音、声音事件),(3) 不同的场景跨度。然而现有数据集在至少一个维度上存在不足,限制了严格而全面的评估。为此,我们提出 JointAVBench,一个具有严格音视频关联性的新基准,涵盖五个认知维度、四种音频类型(语音、声音事件、音乐、声线特征)和三种场景跨度(单场景、跨场景、全场景)。鉴于人工标注成本高昂,我们设计了一套自动化流水线,利用先进的视觉大语言模型、音频大语言模型和通用大语言模型生成严格要求视听联合理解的问答。我们在该数据集上评估了领先的纯视觉、纯音频及多模态模型。结果显示,即使表现最好的 Omni-LLM 平均准确率也仅为 65.3%,虽优于单模态基线,但在跨场景推理方面仍有巨大提升空间。
原文摘要 · Abstract (English)
Understanding videos inherently requires reasoning over both visual and auditory information. To properly evaluate Omni-Large Language Models (Omni-LLMs), which are capable of processing multi-modal information including vision and audio, an effective benchmark must comprehensively cover three key aspects: (1) multi-modal dependency (i.e., questions that cannot be answered using vision or audio alone), (2) diverse audio information types (e.g., speech, sound events), and (3) varying scene spans. However, existing datasets fall short in one or more of these dimensions, limiting strict and comprehensive evaluation. To address this gap, we introduce JointAVBench, a novel benchmark with strict audio-video correlation, spanning five cognitive dimensions, four audio information types (speech, sound events, music, vocal traits), and three scene spans (single-, cross-, and full-scene). Given the high cost of manual annotation, we propose an automated pipeline that leverages state-of-the-art vision-LLMs, audio-LLMs, and general-purpose LLMs to synthesize questions and answers that strictly require joint audio-visual understanding. We evaluate leading vision-only, audio-only, and Omni-LLMs on our dataset. Results show that even the best-performing Omni-LLM achieves an average accuracy of only 65.3\%, outperforming uni-modal baselines but revealing substantial room for improvement, especially in cross-scene reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。