arXiv:2508.14039cs.CV2025-08ICCV被引 8

通过密集修改描述实现更精准的视频检索。

Beyond Simple Edits: Composed Video Retrieval with Dense Modifications

  • 用跨注意力融合视觉与文本信息,对齐细粒度修改描述
  • 在160万样本数据集上达到71.3% Recall@1,提升3.4%
  • 适合需要精确视频内容编辑与检索的研究者

组合式视频检索旨在根据查询视频和包含具体修改描述的文本,检索目标视频。现有框架在细粒度组合查询与时间理解变化方面表现受限。为此,我们提出新数据集Dense-WebVid-CoVR,包含160万条样本,密集修改文本数量是现有数据集的约7倍。我们还设计了基于定位文本编码器的交叉注意力融合模型,实现视觉与文本信息的精准对齐。该模型在所有指标上均达到领先水平,尤其在视觉+文本设置下取得71.3% Recall@1,优于当前最优方法3.4%,验证了其在利用详细描述与密集修改文本方面的有效性。代码与数据集已开源。

原文摘要 · Abstract (English)

Composed video retrieval is a challenging task that strives to retrieve a target video based on a query video and a textual description detailing specific modifications. Standard retrieval frameworks typically struggle to handle the complexity of fine-grained compositional queries and variations in temporal understanding limiting their retrieval ability in the fine-grained setting. To address this issue, we introduce a novel dataset that captures both fine-grained and composed actions across diverse video segments, enabling more detailed compositional changes in retrieved video content. The proposed dataset, named Dense-WebVid-CoVR, consists of 1.6 million samples with dense modification text that is around seven times more than its existing counterpart. We further develop a new model that integrates visual and textual information through Cross-Attention (CA) fusion using grounded text encoder, enabling precise alignment between dense query modifications and target videos. The proposed model achieves state-of-the-art results surpassing existing methods on all metrics. Notably, it achieves 71.3\% Recall@1 in visual+text setting and outperforms the state-of-the-art by 3.4\%, highlighting its efficacy in terms of leveraging detailed video descriptions and dense modification texts. Our proposed dataset, code, and model are available at :https://github.com/OmkarThawakar/BSE-CoVR

视频检索细粒度编辑多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。