arXiv:2503.21699cs.MMcs.AI2025-03AAAI被引 5

构建首个评估多模态大模型音视频理解能力的基准测试

MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX

  • 设计需深度融合音视频信息的2556道题目,覆盖700个视频
  • 顶尖模型最高准确率64%,人类专家达92.8%,差距显著
  • 适合研究音视频联合理解、多模态模型评测的学者使用

我们提出MAVERIX(多模态音视频评估与识别索引),一个统一的基准测试,用于评估多模态大模型在视频理解方面的能力,包含视频、音频和文本输入,并提供人类表现基线。尽管近期具备视觉与音频理解能力的模型已取得显著进展,但该领域仍缺乏标准化的评估框架来全面衡量其跨模态理解性能。MAVERIX从700个视频中收集了2,556道题目,采用多项选择和开放问答形式,专门设计用于检验模型对音视频信息紧密融合的理解能力,涵盖广泛的代理行为场景。该基准首次以细粒度方式系统评估模型在音视频联合感知方面的综合能力。对Qwen 2.5 Omni和Gemini 2.5 Flash-Lite等前沿模型的实验表明,其准确率约为64%,而人类专家达到92.8%的接近天花板水平,揭示了与人类水平之间显著的差距。通过标准化评估协议、严格标注流程和公开工具包,MAVERIX为推动音视频多模态智能发展提供了极具挑战性的测试平台。

原文摘要 · Abstract (English)

We introduce MAVERIX (Multimodal audiovisual Evaluation and Recognition IndeX), a unified benchmark to probe the video understanding in multimodal LLMs, encompassing video, audio, text inputs with human performance baselines. Although recent advancements in models with vision and audio understanding capabilities have shown substantial progress, the field lacks a standardized evaluation framework to thoroughly assess their cross-modality comprehension performance. MAVERIX curates 2,556 questions from 700 videos, in the form of both multiple-choice and open-ended formats, explicitly designed to evaluate multimodal models through questions that necessitate tight integration of video and audio information, spanning a broad spectrum of agentic scenarios. MAVERIX uniquely provides models with audiovisual questions, closely mimicking the multimodal perceptual experiences available to humans during inference and decision-making processes. To our knowledge, MAVERIX is the first benchmark aimed explicitly at assessing comprehensive audiovisual integration in such granularity. Experiments with state-of-the-art models, including Qwen 2.5 Omni and Gemini 2.5 Flash-Lite, show performance around 64% accuracy, while human experts reach near-ceiling performance of 92.8%, exposing a substantial gap to human-level comprehension. With standardized evaluation protocols, a rigorously annotated pipeline, and a public toolkit, MAVERIX establishes a challenging testbed for advancing audiovisual multimodal intelligence.

多模态音视频理解评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。