提出多时刻视频定位新数据集与方法,解决单时刻模型无法应对真实场景问题。
When One Moment Isn't Enough: Multi-Moment Retrieval with Cross-Moment Interactions
- 设计跨时刻交互的FlashMMR框架,通过后验证模块精修时刻边界。
- 在QV-M²数据集上,相比最先进方法提升3.00% G-mAP和2.56% mR@3。
- 适合关注视频时序定位、多时刻理解的研究者与应用开发者。
现有视频时刻定位(MR)方法主要针对单时刻检索(SMR),但实际应用中一个查询可能对应多个相关时刻。为此,本文构建高质量数据集QVHighlights Multi-Moment Dataset(QV-M²),包含2,212条标注覆盖6,384个视频片段,并设计适配多时刻检索(MMR)的新评估指标。在此基础上,提出FlashMMR框架,引入多时刻后验证模块,结合受限时间调整与再评估机制,有效过滤低置信度候选段,实现鲁棒的多时刻对齐。在QV-M²和QVHighlights上对6种现有MR方法进行重训练与评估,结果表明:QV-M²是有效的训练与评测基准,FlashMMR为强基线。在QV-M²上,该方法相较先前最优方法提升3.00% G-mAP、2.70% mAP@3+tgt、2.56% mR@3。所提基准与方法为更真实、更具挑战性的视频时序定位研究奠定基础。代码已开源。
原文摘要 · Abstract (English)
Existing Moment retrieval (MR) methods focus on Single-Moment Retrieval (SMR). However, one query can correspond to multiple relevant moments in real-world applications. This makes the existing datasets and methods insufficient for video temporal grounding. By revisiting the gap between current MR tasks and real-world applications, we introduce a high-quality datasets called QVHighlights Multi-Moment Dataset (QV-M$^2$), along with new evaluation metrics tailored for multi-moment retrieval (MMR). QV-M$^2$ consists of 2,212 annotations covering 6,384 video segments. Building on existing efforts in MMR, we propose a framework called FlashMMR. Specifically, we propose a Multi-moment Post-verification module to refine the moment boundaries. We introduce constrained temporal adjustment and subsequently leverage a verification module to re-evaluate the candidate segments. Through this sophisticated filtering pipeline, low-confidence proposals are pruned, and robust multi-moment alignment is achieved. We retrain and evaluate 6 existing MR methods on QV-M$^2$ and QVHighlights under both SMR and MMR settings. Results show that QV-M$^2$ serves as an effective benchmark for training and evaluating MMR models, while FlashMMR provides a strong baseline. Specifically, on QV-M$^2$, it achieves improvements over prior SOTA method by 3.00% on G-mAP, 2.70% on mAP@3+tgt, and 2.56% on mR@3. The proposed benchmark and method establish a foundation for advancing research in more realistic and challenging video temporal grounding scenarios. Code is released at https://github.com/Zhuo-Cao/QV-M2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。