用文字描述找视频片段,让大模型学会想象和精准定位。
ShotFinder: Imagination-Driven Open-Domain Video Shot Retrieval via Web Search
- 通过文本生成想象视频,再用搜索匹配候选片段。
- 在1210个样本上测试,颜色与风格定位仍难突破。
- 适合研究视频理解、多模态生成的开发者参考。
近年来,大语言模型在信息检索中进展迅速,但主要集中在文本或静态多模态场景。开放域视频片段检索因具有更复杂的时序结构与语义,尚缺乏系统性基准与分析。为此,我们提出ShotFinder,将剪辑需求形式化为以关键帧为导向的片段描述,并引入五类可控单因素约束:时间顺序、色彩、视觉风格、音频和分辨率。从YouTube收集20个主题类别,共1,210个高质量样本,利用大模型生成并经人工验证。基于该基准,我们构建ShotFinder——一个文本驱动的三阶段检索与定位流程:(1) 通过视频想象扩展查询;(2) 利用搜索引擎召回候选视频;(3) 基于描述引导的时间定位。在多个闭源与开源模型上的实验表明,当前性能与人类水平仍有显著差距,且各类约束表现不均:时间定位相对容易,而色彩与视觉风格仍是主要挑战。结果揭示,开放域视频片段检索仍是多模态大模型亟待突破的关键能力。
原文摘要 · Abstract (English)
In recent years, large language models (LLMs) have made rapid progress in information retrieval, yet existing research has mainly focused on text or static multimodal settings. Open-domain video shot retrieval, which involves richer temporal structure and more complex semantics, still lacks systematic benchmarks and analysis. To fill this gap, we introduce ShotFinder, a benchmark that formalizes editing requirements as keyframe-oriented shot descriptions and introduces five types of controllable single-factor constraints: Temporal order, Color, Visual style, Audio, and Resolution. We curate 1,210 high-quality samples from YouTube across 20 thematic categories, using large models for generation with human verification. Based on the benchmark, we propose ShotFinder, a text-driven three-stage retrieval and localization pipeline: (1) query expansion via video imagination, (2) candidate video retrieval with a search engine, and (3) description-guided temporal localization. Experiments on multiple closed-source and open-source models reveal a significant gap to human performance, with clear imbalance across constraints: temporal localization is relatively tractable, while color and visual style remain major challenges. These results reveal that open-domain video shot retrieval is still a critical capability that multimodal large models have yet to overcome.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。