提出可处理任意自然语言查询的视频时序定位框架,突破封闭集限制。
Grounding-MD: Grounded Video-language Pre-training for Open-World Moment Detection
- 通过结构化提示机制支持开放式语言查询,灵活适应多类任务。
- 在4个基准数据集上实现零样本与有监督场景的最先进性能。
- 适合需要开放世界视频理解的应用,如智能监控、内容检索。
时序动作检测与时刻检索是视频理解中的关键任务,旨在精确定位对应特定动作或事件的时间片段。近期研究提出了统一任务的时刻检测方法,但现有方法仍局限于封闭集场景,难以应用于开放世界。为此,我们提出Grounding-MD,一种专为开放世界时刻检测设计的视觉-语言预训练框架。该框架通过结构化提示机制支持任意数量的开放式自然语言查询,实现灵活可扩展的时刻定位。Grounding-MD采用跨模态融合编码器与文本引导融合解码器,促进视频与文本的全面对齐,并支持跨任务协作。在大规模时序动作检测与时刻检索数据集上进行预训练后,模型展现出卓越的语义表征能力,有效应对多样复杂的查询条件。在ActivityNet、THUMOS14、ActivityNet-Captions和Charades-STA四个基准数据集上的综合评估表明,Grounding-MD在零样本与有监督设置下均达到开放世界时刻检测的新最优表现。所有源代码与训练模型将公开发布。
原文摘要 · Abstract (English)
Temporal Action Detection and Moment Retrieval constitute two pivotal tasks in video understanding, focusing on precisely localizing temporal segments corresponding to specific actions or events. Recent advancements introduced Moment Detection to unify these two tasks, yet existing approaches remain confined to closed-set scenarios, limiting their applicability in open-world contexts. To bridge this gap, we present Grounding-MD, an innovative, grounded video-language pre-training framework tailored for open-world moment detection. Our framework incorporates an arbitrary number of open-ended natural language queries through a structured prompt mechanism, enabling flexible and scalable moment detection. Grounding-MD leverages a Cross-Modality Fusion Encoder and a Text-Guided Fusion Decoder to facilitate comprehensive video-text alignment and enable effective cross-task collaboration. Through large-scale pre-training on temporal action detection and moment retrieval datasets, Grounding-MD demonstrates exceptional semantic representation learning capabilities, effectively handling diverse and complex query conditions. Comprehensive evaluations across four benchmark datasets including ActivityNet, THUMOS14, ActivityNet-Captions, and Charades-STA demonstrate that Grounding-MD establishes new state-of-the-art performance in zero-shot and supervised settings in open-world moment detection scenarios. All source code and trained models will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。