用生成视频辅助定位多动词视频片段,提升精准度。
GenSpan: Generation-Calibrated Motion Span Priors for Multi-Verb Video Corpus Moment Retrieval
- 用大模型选字幕线索生成短视频作时间先验
- 在TVR和ActivityNet上提升多动作查询准确率
- 适合需要精细时间定位的视频理解任务
视频语料库片段检索(VCMR)旨在根据自然语言查询定位正确视频及其对应的时间片段,尤其在涉及多个动作的查询中,动作时序关系至关重要。现有方法通常仅依赖文本或静态图像,难以捕捉隐含的动作动态,导致检索错误与时间错位。我们提出GenSpan,一种生成校准的VCMR框架:利用大语言模型筛选字幕线索并分解子事件,生成短辅助视频作为时间先验而非直接检索目标。通过令牌选择器筛选与生成动作对齐的候选视频特征,并采用双向状态空间模型高效预测视频-片段组合。在TVR和ActivityNet-Captions数据集上的实验表明,GenSpan在整体检索和片段定位上均有提升,尤其在复杂多动作查询中表现优异,且相比先进多模态基线显著降低计算成本。
原文摘要 · Abstract (English)
Video Corpus Moment Retrieval (VCMR) aims to retrieve both the correct video and its temporal segment corresponding to a natural-language query, a task that is especially challenging for multi-verb queries where temporal action ordering is critical. Existing approaches often rely solely on text or static images and struggle to capture implicit motion dynamics, leading to retrieval errors and temporal misalignment. We propose GenSpan, a generation-calibrated VCMR framework that constructs short auxiliary videos from LLM-selected subtitle cues and decomposed sub-events, using these as temporal priors rather than direct retrieval targets. A token selector filters candidate-video features aligned with generated motion, and a bidirectional state-space model efficiently predicts video-moment tuples. Experiments on TVR and ActivityNet-Captions demonstrate that GenSpan improves corpus-level retrieval and moment localization, particularly for complex multi-action queries, while reducing computational cost compared to state-of-the-art multimodal baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。