arXiv:2503.17415cs.CVcs.AI2025-03

用视觉语言模型提升视频检索的精准与自适应能力

Enhancing Subsequent Video Retrieval via Vision-Language Models (VLMs)

  • 结合向量搜索与图结构,利用VLM嵌入进行初始检索
  • 通过建模视频片段间上下文关系,实现查询动态优化
  • 适合需要实时交互的视频系统,兼顾精度与扩展性

视频内容的快速增长对高效精准的检索系统提出需求。尽管视觉语言模型(VLM)在表征学习方面表现优异,但在自适应、时间敏感的视频检索中仍存在挑战。本文提出一种新框架,融合向量相似性搜索与基于图的数据结构。利用VLM嵌入进行初始检索,并建模视频片段间的上下文关系,实现查询的自适应优化,提升检索准确率。实验验证了该方法在精度、可扩展性与鲁棒性方面的优势,为动态环境下的交互式视频检索提供了有效解决方案。

原文摘要 · Abstract (English)

The rapid growth of video content demands efficient and precise retrieval systems. While vision-language models (VLMs) excel in representation learning, they often struggle with adaptive, time-sensitive video retrieval. This paper introduces a novel framework that combines vector similarity search with graph-based data structures. By leveraging VLM embeddings for initial retrieval and modeling contextual relationships among video segments, our approach enables adaptive query refinement and improves retrieval accuracy. Experiments demonstrate its precision, scalability, and robustness, offering an effective solution for interactive video retrieval in dynamic environments.

视频检索视觉语言模型图结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。