arXiv:2604.27591cs.CVcs.AI2026-04

通过片段对关系建模,提升视频片段检索的边界预测精度

ClipTBP: Clip-Pair based Temporal Boundary Prediction with Boundary-Aware Learning for Moment Retrieval

论文配图:ClipTBP: Clip-Pair based Temporal Boundary Prediction with Boundary-Aware Learning for Moment Retrieval
图 1 · 摘自论文原文
  • 基于片段对设计边界感知学习机制,显式建模匹配段落间语义关系
  • 引入主边界损失与辅助边界损失,实现更精准的时序边界预测
  • 在模糊查询场景下表现更鲁棒,适配多种现有模型增强性能

视频片段检索任务旨在定位与给定文本查询对应的视频特定时段。现有方法虽通过片段级视觉-语言相似性学习及Transformer时序边界回归提升多模态对齐效果,但未考虑单个查询对应多个答案片段之间的关联关系,易受上下文视觉相似片段干扰,难以排除无关片段。为此,本文提出基于片段对的时序边界预测框架ClipTBP,引入片段级对齐损失以显式学习答案片段间的语义关系,并结合主边界损失与辅助边界损失进行精确时序边界预测。ClipTBP可稳定提升多种现有模型性能,在模糊查询场景下亦表现出更强的边界预测鲁棒性。

原文摘要 · Abstract (English)

Video moment retrieval is the task of retrieving specific segments of a video corresponding to a given text query. Recent studies have been conducted to improve multimodal alignment performance through visual-linguistic similarity learning at the snippet-level and transformer-based temporal boundary regression. However, existing models do not calculate similarity by considering the relationships between multiple answer segments that match the query. Therefore, existing models are easily influenced by visually similar segments in the surrounding context. Existing models calculate similarity at the snippet-level and ignore the relationships between multiple answer segments corresponding to a single query. Therefore, they struggle to exclude segments irrelevant to the query. To address this issues, we propose ClipTBP, a clip-pair temporal boundary prediction framework based on boundary-aware learning. ClipTBP introduces a clip-level alignment loss for explicitly learning the semantic relationship between answer segments. ClipTBP also predicts accurate temporal boundaries by applying both main boundary loss and auxiliary boundary loss. ClipTBP consistently improves performance when applied to various existing models and demonstrates more robust boundary prediction performance even in ambiguous query scenarios.

视频检索边界预测多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。