无需配对数据和标注,联合检索与定位视频段落。
Dual-task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding
- 双任务互增强:检索与定位协同优化,共享特征表示。
- 仅靠粗粒度对齐即可实现精确匹配,减少对标注数据依赖。
- 适合无对应关系、无时间标签的视频理解场景。
视频段落定位(VPG)旨在精准定位与给定文本段落最相关的视频时刻。现有方法通常依赖大规模带时序标注的数据,并假设视频与段落之间已知对应关系,这在实际应用中不现实——构建时序标注成本高,且对应关系常未知。为此,我们提出双任务互增强的联合视频段落检索与定位方法(DMR-JRG)。该方法包含检索分支与定位分支:检索分支通过视频间对比学习,粗略对齐段落与视频的全局特征,缓解模态差异,建立粗粒度特征空间,摆脱对视频-段落对应关系的依赖;该空间进一步辅助定位分支提取细粒度上下文表示。定位分支则通过探索视频片段与文本段落在局部、全局与时间维度上的一致性,实现精准跨模态匹配与定位。通过多维度协同建模,构建细粒度特征空间,显著降低对大规模标注时序标签的需求。
原文摘要 · Abstract (English)
Video Paragraph Grounding (VPG) aims to precisely locate the most appropriate moments within a video that are relevant to a given textual paragraph query. However, existing methods typically rely on large-scale annotated temporal labels and assume that the correspondence between videos and paragraphs is known. This is impractical in real-world applications, as constructing temporal labels requires significant labor costs, and the correspondence is often unknown. To address this issue, we propose a Dual-task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding method (DMR-JRG). In this method, retrieval and grounding tasks are mutually reinforced rather than being treated as separate issues. DMR-JRG mainly consists of two branches: a retrieval branch and a grounding branch. The retrieval branch uses inter-video contrastive learning to roughly align the global features of paragraphs and videos, reducing modality differences and constructing a coarse-grained feature space to break free from the need for correspondence between paragraphs and videos. Additionally, this coarse-grained feature space further facilitates the grounding branch in extracting fine-grained contextual representations. In the grounding branch, we achieve precise cross-modal matching and grounding by exploring the consistency between local, global, and temporal dimensions of video segments and textual paragraphs. By synergizing these dimensions, we construct a fine-grained feature space for video and textual features, greatly reducing the need for large-scale annotated temporal labels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。