构建首个面向长视频关键片段定位的综合性评测基准。
MomentSeeker: A Task-Oriented Benchmark For Long-Video Moment Retrieval
- 基于平均时长超1200秒的多领域长视频,覆盖三类场景任务。
- 支持文本、图像、视频三种查询形式,评估精度与效率挑战显著。
- 适合研究长视频理解、检索模型与多模态交互的学者使用。
准确识别长视频中的关键片段对解决长视频理解(LVU)任务至关重要。然而,现有基准在视频长度和任务多样性上严重受限,或仅关注端到端的LVU性能,难以评估关键片段的精准定位能力。为此,我们提出MomentSeeker,一个新型长视频片段检索(LMVR)评测基准,具备三大特点:第一,基于平均时长达1200秒以上、涵盖电影、异常检测、第一人称视角和体育等多领域的多样化长视频;第二,覆盖全局、事件、对象三个层级的真实场景任务,包括动作识别、目标定位、因果推理等;第三,支持文本仅查询、图像条件查询和视频条件查询等多种形式。在此基础上,我们对生成式方法(直接使用多模态大模型)和检索式方法(借助视频检索器)进行了全面实验。结果表明,尽管最新长视频多模态大模型和任务特定微调有所提升,但在准确性和效率方面仍面临严峻挑战。我们已公开发布MomentSeeker(https://yhy-2000.github.io/MomentSeeker/),以推动该领域研究发展。
原文摘要 · Abstract (English)
Accurately locating key moments within long videos is crucial for solving long video understanding (LVU) tasks. However, existing benchmarks are either severely limited in terms of video length and task diversity, or they focus solely on the end-to-end LVU performance, making them inappropriate for evaluating whether key moments can be accurately accessed. To address this challenge, we propose MomentSeeker, a novel benchmark for long-video moment retrieval (LMVR), distinguished by the following features. First, it is created based on long and diverse videos, averaging over 1200 seconds in duration and collected from various domains, e.g., movie, anomaly, egocentric, and sports. Second, it covers a variety of real-world scenarios in three levels: global-level, event-level, object-level, covering common tasks like action recognition, object localization, and causal reasoning, etc. Third, it incorporates rich forms of queries, including text-only queries, image-conditioned queries, and video-conditioned queries. On top of MomentSeeker, we conduct comprehensive experiments for both generation-based approaches (directly using MLLMs) and retrieval-based approaches (leveraging video retrievers). Our results reveal the significant challenges in long-video moment retrieval in terms of accuracy and efficiency, despite improvements from the latest long-video MLLMs and task-specific fine-tuning. We have publicly released MomentSeeker(https://yhy-2000.github.io/MomentSeeker/) to facilitate future research in this area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。