评测大模型在长视频中多模态联合推理能力,发现现有模型表现远未达标。
MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos
- 构建2万道题的多模态长视频评测集,覆盖13类综合理解任务。
- 顶尖闭源模型准确率仅64.2%,开源模型最高46.8%,差距显著。
- 揭示模型在长视频中跨模态、跨时间推理的系统性失败原因。
多模态大语言模型在孤立的视觉与音频理解任务中表现优异,但在长且复杂的现实世界视频中联合推理多模态(视觉、音频、文本)信号的能力仍缺乏系统评估。我们提出MMOU,一个全新的基准,用于在真实复杂条件下系统评测多模态理解与推理能力。MMOU包含20,000道精心设计的问题,对应11,877段来自网络的视频,时长不一,涵盖多种领域,具有丰富且紧密耦合的音视频内容。该基准覆盖13个基础能力类别,均需跨模态与跨时间整合证据。所有问题均由专业标注员进行多轮人工标注,确保高质量与推理一致性。我们在MMOU上评估了20多个业界领先的大模型。结果表明存在显著性能差距:最佳闭源模型准确率为64.2%,最强开源模型仅为46.8%。研究揭示当前模型在长视频中难以应用基本推理技能,通过深入分析识别出系统性失败模式,并为改进方向提供洞见。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have shown strong performance in visual and audio understanding when evaluated in isolation. However, their ability to jointly reason over omni-modal (visual, audio, and textual) signals in long and complex videos remains largely unexplored. We introduce MMOU, a new benchmark designed to systematically evaluate multimodal understanding and reasoning under these challenging, real-world conditions. MMOU consists of 20,000 carefully curated questions paired with 11877 web-collected videos of varying length, spanning diverse domains and exhibiting rich, tightly coupled audio-visual content. The benchmark covers 13 fundamental skill categories, all of which require integrating evidence across modalities and time. All questions are manually annotated across multiple turns by professional annotators, ensuring high quality and reasoning fidelity. We evaluate 20+ state-of-the-art open-source and proprietary multimodal models on MMOU. The results expose substantial performance gaps: the best closed-source model achieves only 64.2% accuracy, while the strongest open-source model reaches just 46.8%. Our results highlight the challenges of long-form omni-modal understanding, revealing that current models frequently fail to apply even fundamental skills in long videos. Through detailed analysis, we further identify systematic failure modes and provide insights into where and why current models break.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。