arXiv:2504.13035cs.CVcs.AI2025-04ICCV被引 5

用原型压缩视频上下文,兼顾检索精度与效率

Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval

  • 将视频多尺度上下文压缩为固定数量原型,降低计算开销
  • 在TVR等数据集上达到领先精度,且推理速度更快
  • 适合需要高效精准视频检索的应用场景

在检索系统中,同时实现搜索精度与效率具有根本性挑战。这一挑战在部分相关视频检索(PRVR)中尤为突出:为提升精度,需引入多样化的视频上下文表示,但会显著增加计算与内存开销。为此,我们提出一种原型化PRVR框架,将视频内多样化上下文编码为固定数量的原型,并设计多种策略增强原型中的文本关联与视频理解,同时引入正交目标确保原型覆盖多样内容。为使原型可被文本查询检索,同时准确编码视频上下文,我们实施跨模态与单模态重建任务:前者在共享空间中对齐原型与文本特征,后者在编码过程中保留全部视频上下文。此外,采用视频混合技术提供弱监督以进一步对齐原型与文本表示。在TVR、ActivityNet-Captions和QVHighlights上的大量实验验证了该方法的有效性,且未牺牲效率。

原文摘要 · Abstract (English)

In a retrieval system, simultaneously achieving search accuracy and efficiency is inherently challenging. This challenge is particularly pronounced in partially relevant video retrieval (PRVR), where incorporating more diverse context representations at varying temporal scales for each video enhances accuracy but increases computational and memory costs. To address this dichotomy, we propose a prototypical PRVR framework that encodes diverse contexts within a video into a fixed number of prototypes. We then introduce several strategies to enhance text association and video understanding within the prototypes, along with an orthogonal objective to ensure that the prototypes capture a diverse range of content. To keep the prototypes searchable via text queries while accurately encoding video contexts, we implement cross- and uni-modal reconstruction tasks. The cross-modal reconstruction task aligns the prototypes with textual features within a shared space, while the uni-modal reconstruction task preserves all video contexts during encoding. Additionally, we employ a video mixing technique to provide weak guidance to further align prototypes and associated textual representations. Extensive evaluations on TVR, ActivityNet-Captions, and QVHighlights validate the effectiveness of our approach without sacrificing efficiency.

视频检索原型学习多尺度上下文高效检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。