构建真实视频搜索基准与智能代理框架,解决模糊记忆下的视频定位难题。
Beyond Closed-Pool Video Retrieval: A Benchmark and Agent Framework for Real-World Video Search and Moment Localization

- 提出多维描述框架,模拟人类真实记忆的模糊搜索线索。
- 构建含1440样本的开源基准,覆盖20类场景与4种时长分布。
- 设计推理代理系统,实现类似人类回忆-搜索-验证的认知流程。
传统视频检索评测聚焦精确描述与封闭视频池匹配,无法反映开放网络中基于模糊、多维度记忆的真实搜索。本文提出RVMS-Bench,一个全面评估真实世界视频记忆搜索的基准系统,包含1440个样本,覆盖20个多样化类别和4个时长组,均来自真实开放网络视频。该基准采用分层描述框架,涵盖全局印象、关键瞬间、时间上下文与听觉记忆,完整模拟多维度搜索线索,并通过人机协同验证确保质量。进一步提出RACLO智能体框架,利用溯因推理模拟人类‘回忆-搜索-验证’认知过程,有效应对真实世界中基于模糊记忆的视频检索挑战。实验表明,现有多模态大模型在模糊记忆驱动的视频检索与片段定位任务中仍表现不足。本工作有望推动视频检索在现实非结构化场景中的鲁棒性发展。
原文摘要 · Abstract (English)
Traditional video retrieval benchmarks focus on matching precise descriptions to closed video pools, failing to reflect real-world searches characterized by fuzzy, multi-dimensional memories on the open web. We present \textbf{RVMS-Bench}, a comprehensive system for evaluating real-world video memory search. It consists of \textbf{1,440 samples} spanning \textbf{20 diverse categories} and \textbf{four duration groups}, sourced from \textbf{real-world open-web videos}. RVMS-Bench utilizes a hierarchical description framework encompassing \textbf{Global Impression, Key Moment, Temporal Context, and Auditory Memory} to mimic realistic multi-dimensional search cues, with all samples strictly verified via a human-in-the-loop protocol. We further propose \textbf{RACLO}, an agentic framework that employs abductive reasoning to simulate the human ``Recall-Search-Verify'' cognitive process, effectively addressing the challenge of searching for videos via fuzzy memories in the real world. Experiments reveal that existing MLLMs still demonstrate insufficient capabilities in real-world Video Retrieval and Moment Localization based on fuzzy memories. We believe this work will facilitate the advancement of video retrieval robustness in real-world unstructured scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。