arXiv:2609.08999cs.CV2026-09

解决视频检索中局部匹配误导问题,通过条件验证提升结果可靠性

Concentrate After Imagination: Text-Conditioned Evidence Grounding for Partially Relevant Video Retrieval

论文配图:Concentrate After Imagination: Text-Conditioned Evidence Grounding for Partially Relevant Video Retrieval
图 1 · 摘自论文原文
  • 设计得分级证据验证机制,动态筛选与查询相关的视频片段
  • 在三个基准上均达最优,最高提升1.5分,显著改善检索精度
  • 适合关注细粒度视频理解与检索的科研人员和工程师

部分相关视频检索(PRVR)旨在查询仅描述短时段时,从非剪辑视频中定位相关内容。尽管现有方法提升了局部表征、不确定性建模和全局上下文,最终排序仍常依赖最强局部响应,导致偶然相似片段产生错误峰值。本文识别此为查询无关的集中瓶颈,提出TRACE——一种针对PRVR的得分级证据验证算子。给定查询与全局视频注册信息,TRACE激活与查询相关的注册项,将支持传递至帧级证据,并在局部时间选择前平滑地对替代的查询-注册-帧路径进行边际化。不同于表示层特征融合,TRACE仅以查询条件化的残差校准原始局部得分。在ActivityNet Captions、Charades-STA和TVR数据集上,TRACE在所有三者上均取得最佳SumR表现,分别使DreamPRVR骨干网络提升1.2、1.1和1.5分。消融实验、路由干扰、难负样本及跨骨干迁移分析表明,性能提升源于查询条件化的证据验证,而非通用得分偏移。

原文摘要 · Abstract (English)

Partially Relevant Video Retrieval (PRVR) retrieves untrimmed videos when queries describe only short moments. Although recent methods improve local representations, uncertainty modeling, and global context, final ranking often still trusts the strongest local response; a coincidentally similar fragment can therefore produce an unsupported peak. We identify this failure as the query-agnostic concentration bottleneck and propose TRACE, a score-level evidence verification operator for PRVR. Given a query and global video registers, TRACE activates query-relevant registers, routes their support to frame-level evidence, and smoothly marginalizes alternative query-to-register-to-frame paths before localized temporal selection. Unlike representation-level feature fusion, TRACE uses this evidence only as a query-conditioned residual calibration of the original local score. On ActivityNet Captions, Charades-STA, and TVR, TRACE achieves the best SumR on all three benchmarks and improves the DreamPRVR backbone by 1.2, 1.1, and 1.5 points, respectively. Ablation, routing-corruption, hard-negative, and cross-backbone transfer analyses support the interpretation that the gains arise from query-conditioned evidence verification rather than a generic score offset.

视频检索证据验证多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。