arXiv:2508.06869cs.CVcs.AI2025-08中稿 · CVPR被引 2

融合视觉与字幕信息,精准挑选长视频关键帧

VSI: Visual Subtitle Integration for Keyframe Selection to enhance Long Video Understanding

  • 双分支协同检索:结合视频搜索与字幕匹配
  • 在长视频理解任务中准确率领先,文本相关任务提升显著
  • 适合需要精准语义定位的长视频分析场景

多模态大语言模型(MLLM)在视觉-语言任务中表现优异,但处理长视频时受限于输入上下文长度和高计算成本。因此稀疏帧采样成为必要预处理步骤,采样帧的质量直接影响下游性能。现有关键帧搜索算法虽兼顾效率与质量,但过度依赖视觉模态,难以适应文本相关任务,常导致检索结果偏离核心语义内容。为此,我们提出视觉-字幕融合(VSI)框架,采用双分支协同检索机制,结合视频搜索与字幕匹配,融合互补的视觉与文本信息以实现精确定位。在LongVideoBench和VideoMME上的实验表明,VSI在关键帧检索中达到最先进水平,并在文本相关任务中实现突破性表现,且在其他任务上展现出强泛化能力。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) demonstrate exceptional performance in vision-language tasks, yet their processing of long videos is constrained by input context length and high computational costs. Sparse frame sampling thus becomes a necessary preprocessing step, with sampled frame quality directly impacting downstream performance. Existing keyframe search algorithms achieve a balance between efficiency and sampled frame quality but heavily rely on the visual modality alone. This makes them difficult to adapt to text-related tasks and often leads to retrieval results deviating from core semantic content. To address this, we propose the VISUAL-SUBTITLE INTEGRATION (VSI), a multimodal keyframe retrieval framework. It employs a dual-branch collaborative retrieval approach combining Video Search and Subtitle Match to fuse complementary visual and textual information for precise localization. Experiments on LongVideoBench and VideoMME demonstrate that VSI achieves state-of-the-art accuracy in keyframe retrieval while delivering breakthrough performance in text-related tasks and exhibiting strong generalization across other tasks.

长视频理解关键帧选择多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。