arXiv:2603.23902cs.CVcs.AI2026-03中稿 · ICME 2026

提升视频片段检索精度,解决语义与视觉信息不匹配问题

Knowledge-Refined Dual Context-Aware Network for Partially Relevant Video Retrieval

  • 分层语义聚合增强查询语义表达
  • 动态时间注意力捕捉关键事件与时间连贯性
  • 基于CLIP的动态知识蒸馏提升检索准确性

从非剪辑视频中检索部分相关片段仍面临两大挑战:文本与视频片段间信息密度不匹配,以及注意力机制难以捕捉语义焦点和事件关联。本文提出KDC-Net,一种知识精炼的双上下文感知网络,从文本与视觉双视角应对上述问题。文本侧采用分层语义聚合模块,捕获并自适应融合多尺度短语线索以丰富查询语义;视频侧引入动态时间注意力机制,结合相对位置编码与自适应时间窗口,突出具有局部时间一致性的关键事件。此外,通过增强时间连续性感知的动态CLIP知识蒸馏策略,实现段落感知且目标对齐的知识迁移。在PRVR基准测试中,KDC-Net持续优于现有最优方法,尤其在低片段-视频比条件下表现更优。

原文摘要 · Abstract (English)

Retrieving partially relevant segments from untrimmed videos remains difficult due to two persistent challenges: the mismatch in information density between text and video segments, and limited attention mechanisms that overlook semantic focus and event correlations. We present KDC-Net, a Knowledge-Refined Dual Context-Aware Network that tackles these issues from both textual and visual perspectives. On the text side, a Hierarchical Semantic Aggregation module captures and adaptively fuses multi-scale phrase cues to enrich query semantics. On the video side, a Dynamic Temporal Attention mechanism employs relative positional encoding and adaptive temporal windows to highlight key events with local temporal coherence. Additionally, a dynamic CLIP-based distillation strategy, enhanced with temporal-continuity-aware refinement, ensures segment-aware and objective-aligned knowledge transfer. Experiments on PRVR benchmarks show that KDC-Net consistently outperforms state-of-the-art methods, especially under low moment-to-video ratios.

视频检索注意力机制知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。