arXiv:2510.10689cs.AI2025-10被引 41

构建首个评估音视频协同理解的基准,揭示当前模型真实推理能力短板

OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs

  • 设计1000组带分步推理的音视频QA对,覆盖13类复杂理解任务
  • 模型在该基准上表现远低于人类,开源模型显著落后于闭源模型
  • 专为检测模态互补与逻辑一致性而设,适合评估真正跨模态推理能力

近年来多模态大语言模型在视频理解方面展现出巨大潜力。然而,现有评测基准未能全面评估音频与视觉模态间的协同推理能力,常忽略某一模态或逻辑整合不当。为此,我们提出OmniVideoBench,一个大规模、严谨设计的基准,专注于评估音视频协同理解,强调模态互补性与逻辑一致性。OmniVideoBench包含1000个高质量问答对,每对均附有逐步推理轨迹,源自628段时长从数秒至30分钟的多样化视频,并经人工验证确保正确性与唯一性。该基准涵盖13种精心设计的问题类型,涵盖时间推理、空间定位、计数、因果推断、摘要等核心挑战。对多个MLLM在OmniVideoBench上的评估显示,模型性能与人类推理存在显著差距,开源模型明显落后于闭源模型,凸显真实音视频推理的固有难度。我们将公开OmniVideoBench,以推动具备更强泛化推理能力的MLLM发展。

原文摘要 · Abstract (English)

Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities across audio and visual modalities, often neglecting either one of the modalities or integrating them in a logically inconsistent manner. To bridge this gap, we introduce OmniVideoBench, a large-scale and rigorously designed benchmark dedicated to assessing synergistic audio-visual understanding, with a strong emphasis on modality complementarity and logical consistency. Specifically, OmniVideoBench comprises 1000 high-quality question-answer(QA) pairs, each annotated with step-by-step reasoning traces, derived from 628 diverse videos ranging from several seconds to 30 minutes, and manually verified to guarantee complete correctness and uniqueness. Moreover, OmniVideoBench encompasses 13 carefully designed question types, covering temporal reasoning, spatial localization, counting, causal inference, summarization, and beyond, thereby capturing the essential challenges of video understanding. Evaluation of multiple MLLMs on OmniVideoBench reveals a pronounced gap between model performance and human reasoning, with open-source models lagging significantly behind their closed-source counterparts, underscoring the inherent difficulty of genuine audio-visual reasoning. We will release OmniVideoBench to foster the development of MLLMs with stronger and more generalizable reasoning capabilities.

多模态视频理解评测基准音视频融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。