构建多轮对话视频检索数据集,让系统能像真人一样交互式找视频。
IVCR-200K: A Large-Scale Multi-turn Dialogue Benchmark for Interactive Video Corpus Retrieval
- 提出多轮对话式视频检索新任务,支持用户与系统动态交互。
- 构建20万条双语多轮对话数据,覆盖完整视频与片段检索。
- 基于多模态大模型实现可解释的交互式检索,适合智能搜索研发者。
近年来,视频检索与视频片段检索任务取得显著进展,分别实现根据文本查询返回完整视频或特定片段。这些成果极大提升了用户的搜索满意度。然而,现有方法缺乏系统与用户之间的有效互动,单向检索范式难以满足至少80.8%用户的个性化和动态需求。本文提出交互式视频语料库检索(IVCR)任务,模拟真实场景中用户与系统间的多轮、对话式交互。为推动该任务研究,我们构建了IVCR-200K数据集——一个高质量、双语、多轮、对话式且支持抽象语义理解的大规模数据集,涵盖视频及片段级检索。此外,我们设计了一个基于多模态大语言模型(MLLMs)的综合框架,支持用户以多种模式进行交互,并提供更具可解释性的检索结果。大量实验验证了该数据集与框架的有效性。
原文摘要 · Abstract (English)
In recent years, significant developments have been made in both video retrieval and video moment retrieval tasks, which respectively retrieve complete videos or moments for a given text query. These advancements have greatly improved user satisfaction during the search process. However, previous work has failed to establish meaningful "interaction" between the retrieval system and the user, and its one-way retrieval paradigm can no longer fully meet the personalization and dynamic needs of at least 80.8\% of users. In this paper, we introduce the Interactive Video Corpus Retrieval (IVCR) task, a more realistic setting that enables multi-turn, conversational, and realistic interactions between the user and the retrieval system. To facilitate research on this challenging task, we introduce IVCR-200K, a high-quality, bilingual, multi-turn, conversational, and abstract semantic dataset that supports video retrieval and even moment retrieval. Furthermore, we propose a comprehensive framework based on multi-modal large language models (MLLMs) to help users interact in several modes with more explainable solutions. The extensive experiments demonstrate the effectiveness of our dataset and framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。