arXiv:2605.22907cs.CV2026-05被引 2

构建超长视频理解新基准,考验模型持续推理与多模态感知能力。

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding

论文配图:VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding
图 1 · 摘自论文原文
  • 以连续观看时长为衡量标准,设计跨11个领域54类的超长视频数据集。
  • 平均连续问题需看16分钟视频,最长跨度达小时级,远超现有基准。
  • 支持视觉与音视频联合评估,适合测试大模型在复杂场景下的长期理解力。

真实世界中的长视频理解需要模型在极长时间跨度内实现持续跟踪、信息整合与记忆保留。当前基准虽已扩展视频时长,但多数任务仅考察短片段理解,难以体现超长上下文推理的真实挑战。为此,我们提出VideoOdyssey,一个专为超长上下文与多模态视频理解设计的基准。其特点包括:1)极端视频时长与多样性,覆盖11个领域54个子类别,平均时长109分钟;2)全面评估场景,提供VideoOdyssey-V(专注多模态大模型视觉理解极限)与VideoOdyssey-AV(评估音视频同步理解)两个子集;3)超长且多层次的连续证书,平均连续问题长度达16分钟(VideoOdyssey-V)和12.8分钟(VideoOdyssey-AV),涵盖从秒级到小时级共5个粒度层级,可系统诊断模型在不同上下文长度下的表现。大量实验表明,当前多模态大模型的瓶颈不仅限于简单检索,更体现在跨长度持续推理、细粒度感知及非语言多模态理解等方面。

原文摘要 · Abstract (English)

Real-world long video understanding requires models to perform continuous tracking, information integration and memory retention over massive temporal spans within extreme video durations. Mastering this intense cognitive load constitutes the fundamental bottleneck in long video understanding. While existing benchmarks have driven progress by scaling up video duration, their evaluation tasks often require comprehending only short and isolated video segments, falling short of capturing the challenge of ultra-long-context reasoning. To measure this cognitive load, we emphasize continuous certificate length, defined as the video length a human must continuously watch to definitively answer a given question. Driven by this metric, we introduce VideoOdyssey, a benchmark specifically designed for ultra-long-context and omni-modal video understanding. VideoOdyssey is characterized by three key features: 1) Extreme video duration and diversity: spanning 11 domains and 54 subcategories with an average video duration of 109 minutes; 2) Comprehensive evaluation scenarios: offering two subsets to address different research focuses, i.e., VideoOdyssey-V for probing the limits of visual understanding in MLLMs, and VideoOdyssey-AV for evaluating synchronized audio-visual understanding for omni-modal models; 3) Ultra-long and multi-level continuous certificates: extending the average continuous certificate to 16 minutes for VideoOdyssey-V and 12.8 minutes for VideoOdyssey-AV. Crucially, we design 5 granular levels from seconds to hours, providing a comprehensive diagnostic tool to evaluate models across varying context lengths and cognitive loads. Extensive evaluations show that bottlenecks of current MLLMs extend beyond simple retrieval to include struggles with continuous reasoning across varying context lengths, fine-grained perception, and non-verbal omni-modal understanding.

视频理解多模态长上下文评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。