arXiv:2501.05510cs.CVcs.AI2025-01CVPR被引 131

新基准OVO-Bench评估视频大模型实时理解能力,揭示当前模型与人类差距。

OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?

论文配图:OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
图 1 · 摘自论文原文
  • 设计三种时间场景:回溯、实时、前瞻响应,测试模型对时间戳的敏感度
  • 包含644个视频、2800条带时间标记的细粒度标注,兼顾自动化生成与人工校验
  • 9个视频大模型在该基准上表现均低于人类,凸显在线理解能力短板

时间感知是区分离线与在线视频大模型的关键能力。离线模型依赖完整视频进行静态分析,而在线模型需随视频流动态处理并根据提问时间点调整回应。现有基准未充分评估此能力,为此我们提出OVO-Bench(Online-VideO-Benchmark),强调时间戳在高级在线视频理解评测中的重要性。该基准涵盖12项任务,包含644个独特视频和约2800条经人工精心标注的细粒度时间元信息。通过自动化生成与人工校验结合的方式构建高质量样本,并建立系统化的时间轴查询评估流程。对九个视频大模型的测评显示,尽管在传统基准上取得进展,当前模型在在线视频理解方面仍显著落后于人类代理。我们希望OVO-Bench能推动视频大模型发展,激发未来在线视频推理研究。项目代码与数据集见https://github.com/JoeLeelyf/OVO-Bench。

原文摘要 · Abstract (English)

Temporal Awareness, the ability to reason dynamically based on the timestamp when a question is raised, is the key distinction between offline and online video LLMs. Unlike offline models, which rely on complete videos for static, post hoc analysis, online models process video streams incrementally and dynamically adapt their responses based on the timestamp at which the question is posed. Despite its significance, temporal awareness has not been adequately evaluated in existing benchmarks. To fill this gap, we present OVO-Bench (Online-VideO-Benchmark), a novel video benchmark that emphasizes the importance of timestamps for advanced online video understanding capability benchmarking. OVO-Bench evaluates the ability of video LLMs to reason and respond to events occurring at specific timestamps under three distinct scenarios: (1) Backward tracing: trace back to past events to answer the question. (2) Real-time understanding: understand and respond to events as they unfold at the current timestamp. (3) Forward active responding: delay the response until sufficient future information becomes available to answer the question accurately. OVO-Bench comprises 12 tasks, featuring 644 unique videos and approximately human-curated 2,800 fine-grained meta-annotations with precise timestamps. We combine automated generation pipelines with human curation. With these high-quality samples, we further developed an evaluation pipeline to systematically query video LLMs along the video timeline. Evaluations of nine Video-LLMs reveal that, despite advancements on traditional benchmarks, current models struggle with online video understanding, showing a significant gap compared to human agents. We hope OVO-Bench will drive progress in video LLMs and inspire future research in online video reasoning. Our benchmark and code can be accessed at https://github.com/JoeLeelyf/OVO-Bench.

视频理解时间感知大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。