提出通用视频片段检索新任务,支持多段匹配与无匹配情况。
Retrieving Any Relevant Moments: Benchmark and Models for Generalized Moment Retrieval

- 构建可灵活生成的足球视频数据集,支持多段与空匹配场景。
- 引入互补评估指标,全面衡量模型在空集拒绝与定位上的表现。
- 设计轻量适配器与专用奖励机制,显著提升现有模型性能。
视频片段检索(VMR)旨在定位与自然语言查询对应的视频时间片段,但传统方法通常假设每条查询仅对应一个匹配片段。该假设在真实场景中不成立,因查询可能对应多个或零个片段。为此,我们提出通用片段检索(GMR),统一要求检索全部相关片段或预测空集。为系统研究GMR,我们构建了基于复杂足球视频的大规模基准Soccer-GMR,包含真实负样本与正样本。该基准通过可伸缩的半自动化流水线结合人工验证生成,保障高质量标注。我们还设计了涵盖空集拒绝、正样本定位与端到端性能的联合评估协议。最后,我们建立了两种范式下的强基线:一种是适用于判别式VMR模型的轻量级插件式适配器;另一种是用于微调多模态大语言模型(MLLMs)的专属GRPO奖励。大量实验显示,在所有指标上均有持续提升,并揭示当前方法的关键局限,确立GMR作为更真实且更具挑战性的视频-语言理解基准。
原文摘要 · Abstract (English)
Video Moment Retrieval (VMR) aims to localize temporal segments in videos that correspond to a natural language query, but typically assumes only a single matching moment for each query. This assumption does not always hold in real-world scenarios, where queries may correspond to multiple or no moments. Thus, we formulate Generalized Moment Retrieval (GMR), a unified setting that requires retrieving the complete set of relevant moments or predicting an empty set. To enable systematic study of GMR, we introduce Soccer-GMR, a large-scale benchmark built on challenging soccer videos that reflect general GMR scenarios, with realistic negative and positive queries. The benchmark is constructed via a duration-flexible semi-automated pipeline with human verification, enabling scalable data generation while maintaining high annotation quality. We further design a unified evaluation protocol with complementary metrics tailored for null-set rejection, positive-query localization, and end-to-end GMR performance. Finally, we establish strong baselines across two modeling paradigms: a lightweight plug-and-play GMR adapter for discriminative VMR models, and a GMR-tailored GRPO reward for fine-tuning multimodal large language models (MLLMs). Extensive experiments show consistent gains across all metrics and expose key limitations of current methods, positioning GMR as a more realistic and challenging benchmark for video-language understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。