arXiv:2503.19009cs.CVcs.IR2025-03CVPR被引 19

提出Video-ColBERT,用细粒度交互提升文本查视频的准确率

Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval

  • 采用时空令牌级交互,精细匹配文本与视频内容
  • 在多个基准上优于传统双编码器方法,性能显著提升
  • 适合需要高精度图文视频检索的研究与应用

本文研究文本到视频检索(T2VR)问题。受文本-文档、文本-图像及文本-视频检索中后期交互技术成功的启发,我们提出了Video-ColBERT,一种简单高效的细粒度相似性评估机制。该方法基于三个核心组件:细粒度的时空令牌级交互、查询与视觉内容扩展,以及训练中的双Sigmoid损失。我们发现这种交互与训练范式能够生成强且兼容的视频表征。相比其他双编码器方法,该模型在常见文本到视频检索基准测试中表现出更优性能。

原文摘要 · Abstract (English)

In this work, we tackle the problem of text-to-video retrieval (T2VR). Inspired by the success of late interaction techniques in text-document, text-image, and text-video retrieval, our approach, Video-ColBERT, introduces a simple and efficient mechanism for fine-grained similarity assessment between queries and videos. Video-ColBERT is built upon 3 main components: a fine-grained spatial and temporal token-wise interaction, query and visual expansions, and a dual sigmoid loss during training. We find that this interaction and training paradigm leads to strong individual, yet compatible, representations for encoding video content. These representations lead to increases in performance on common text-to-video retrieval benchmarks compared to other bi-encoder methods.

文本到视频检索双编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。