让视频搜索能像对话一样逐步优化,一次交互就达74.3%准确率。
ReCoVR: Closing the Loop in Interactive Composed Video Retrieval

- 双路径架构:一路听指令,一路自检检索过程
- 单轮交互后在WebVid数据集上达到74.30% R@1
- 适合需要反复调整搜索目标的智能视频系统
组合视频检索(CoVR)通过参考视频和修改文本查找目标视频,但现有方法仅支持单轮交互,无法适应真实场景中渐进式视觉搜索的需求。为此,我们首次形式化了交互式组合视频检索——一种多轮扩展的CoVR设置,用户可通过自然语言反馈逐步细化搜索意图。将现有交互检索方法适配于此场景时发现两个结构性缺陷:依赖单一检索通道,以及开环设计导致系统虽接收用户反馈却无法诊断自身检索轨迹是否偏离或停滞。为此,我们提出ReCoVR(反射式组合视频检索),基于反射感知的双路径架构:意图路径将异构反馈分配至互补检索通道;反思路径则对检索轨迹进行层级分析,监控结果演化并纠正多轮中的错误。在多个基准测试上的实验表明,ReCoVR持续优于交互基线,在WebVid-CoVR-Test数据集上仅经一轮交互即实现74.30% R@1。
原文摘要 · Abstract (English)
Composed video retrieval (CoVR) searches for target videos using a reference video and a modification text, but existing methods are restricted to a single interaction round and cannot support the progressive nature of real-world visual search. To bridge this gap, we first formalize interactive composed video retrieval, a multi-turn extension of CoVR, where users progressively refine their search intent through natural-language feedback across turns. Adapting existing interactive retrieval methods to this setting reveals two structural weaknesses: reliance on a single retrieval channel and an open-loop retrieval design that consumes user feedback but does not diagnose whether its own retrieval trajectory is drifting or stagnating. To address these limitations, we propose ReCoVR (Reflexive Composed Video Retrieval), a dual-pathway architecture built on reflexive perception, where the system treats its retrieval history as diagnostic evidence alongside user feedback. Specifically, an Intent Pathway routes heterogeneous feedback to complementary retrieval channels, while a Reflection Pathway performs trajectory-level reflection to monitor result evolution and correct retrieval errors across turns. Experiments on multiple benchmarks show that ReCoVR consistently outperforms interactive baselines, notably achieving 74.30% R@1 after just one interactive round on the WebVid-CoVR-Test dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。