arXiv:2412.15678cs.CV2024-12AAAI被引 32

多对视频句子定位新方法,提升跨模态对齐效率与泛化能力。

Multi-Pair Temporal Sentence Grounding via Multi-Thread Knowledge Transfer Network

  • 通过多线程知识迁移网络协同训练多组视频-句子对。
  • 自监督对比模块增强模态内与跨模态语义一致性。
  • 原型对齐与自适应负样本筛选,提升复杂场景定位精度。

给定未剪辑视频与句子查询的视频-查询对,时序句子定位(TSG)旨在定位视频中与查询相关的片段。尽管现有优秀TSG方法已取得显著成果,但它们通常独立训练每对视频-查询,忽略不同对之间的关联。我们观察到,相似的视频或查询内容不仅能帮助模型更好地理解并泛化跨模态表示,还能辅助定位复杂对。传统方法采用单线程框架,无法协同训练多对,且反复重建冗余知识,限制了实际应用。为此,本文提出全新设置:多对TSG,旨在协同训练多个视频-查询对。我们设计了一种新型视频-查询协同训练方法——多线程知识迁移网络,有效且高效地定位多种视频-查询对。首先,挖掘不同查询间的时空语义以相互协作;为同时学习模态内与跨模态表示,设计了自监督策略的跨模态对比模块,探索语义一致性;为充分对齐不同对间的视觉与文本表示,提出原型对齐策略:1)匹配物体原型与短语原型实现空间对齐;2)对齐动作原型与句子原型实现时间对齐。最后,开发自适应负样本选择模块,动态生成跨模态匹配阈值。大量实验验证了所提方法的有效性与高效性。

原文摘要 · Abstract (English)

Given some video-query pairs with untrimmed videos and sentence queries, temporal sentence grounding (TSG) aims to locate query-relevant segments in these videos. Although previous respectable TSG methods have achieved remarkable success, they train each video-query pair separately and ignore the relationship between different pairs. We observe that the similar video/query content not only helps the TSG model better understand and generalize the cross-modal representation but also assists the model in locating some complex video-query pairs. Previous methods follow a single-thread framework that cannot co-train different pairs and usually spends much time re-obtaining redundant knowledge, limiting their real-world applications. To this end, in this paper, we pose a brand-new setting: Multi-Pair TSG, which aims to co-train these pairs. In particular, we propose a novel video-query co-training approach, Multi-Thread Knowledge Transfer Network, to locate a variety of video-query pairs effectively and efficiently. Firstly, we mine the spatial and temporal semantics across different queries to cooperate with each other. To learn intra- and inter-modal representations simultaneously, we design a cross-modal contrast module to explore the semantic consistency by a self-supervised strategy. To fully align visual and textual representations between different pairs, we design a prototype alignment strategy to 1) match object prototypes and phrase prototypes for spatial alignment, and 2) align activity prototypes and sentence prototypes for temporal alignment. Finally, we develop an adaptive negative selection module to adaptively generate a threshold for cross-modal matching. Extensive experiments show the effectiveness and efficiency of our proposed method.

时序定位跨模态对齐协同训练自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。