构建首个聚焦音频的视频理解评测基准,破解纯文本解题陷阱
Audio-centric Video Understanding Benchmark without Text Shortcut
- 设计以音频为核心的视频理解任务,检验模型对声音内容与音画交互的理解能力
- 提出答案置换过滤机制,有效消除仅靠问题文本即可作答的作弊漏洞
- 覆盖开源与闭源多模型评估,揭示当前音视频大模型在听觉理解上的短板
音频在音视频大模型的视频理解任务中常被视为辅助模态,仅用于补充视觉信息。然而,视频的深度理解高度依赖音频,因其提供关键语境、情感线索和语义意义,这些是视觉数据难以单独呈现的。本文提出一个以音频为中心的视频理解评测基准(AVUT),旨在评估多模态大模型在音频信息理解方面的表现。AVUT包含一系列精心设计的音频核心任务,全面测试模型对音频内容及音视频交互的理解能力。同时,本工作指出现有基准普遍存在文本捷径问题——正确答案可仅通过问题文本推断,无需观看视频。为解决此问题,提出基于答案置换的过滤机制。通过对多种开源与专有模型的全面评估,揭示了当前音视频大模型在音频理解方面的系统性不足。演示与数据集已公开于 https://github.com/lark-png/AVUT。
原文摘要 · Abstract (English)
Audio often serves as an auxiliary modality in video understanding tasks of audio-visual large language models (LLMs), merely assisting in the comprehension of visual information. However, a thorough understanding of videos significantly depends on auditory information, as audio offers critical context, emotional cues, and semantic meaning that visual data alone often lacks. This paper proposes an audio-centric video understanding benchmark (AVUT) to evaluate the video comprehension capabilities of multimodal LLMs with a particular focus on auditory information. AVUT introduces a suite of carefully designed audio-centric tasks, holistically testing the understanding of both audio content and audio-visual interactions in videos. Moreover, this work points out the text shortcut problem that largely exists in other benchmarks where the correct answer can be found from question text alone without needing videos. AVUT addresses this problem by proposing a answer permutation-based filtering mechanism. A thorough evaluation across a diverse range of open-source and proprietary multimodal LLMs is performed, followed by the analyses of deficiencies in audio-visual LLMs. Demos and data are available at https://github.com/lark-png/AVUT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。