arXiv:2510.21406cs.CV2025-10NeurIPS被引 3

构建多模态长视频检索基准,支持细粒度查询与多层级视觉匹配。

MUVR: A Multi-Modal Untrimmed Video Retrieval Benchmark with Multi-Level Visual Correspondence

  • 提出基于核心内容的六级视觉对应机制,覆盖事件、动作等多层匹配需求。
  • 包含5.3万段未剪辑视频和1050个跨模态查询,支持模型性能多维评估。
  • 适合研究长视频检索、多模态理解及大模型重排序能力的学者使用。

我们提出了多模态未剪辑视频检索任务,并构建了新的基准MUVR,以推动长视频平台的视频检索发展。MUVR旨在通过多模态查询检索包含相关片段的未剪辑视频。其特点包括:1)实用的检索范式:支持以视频为中心的多模态查询,通过长文本描述、视频标签提示和掩码提示表达细粒度检索需求;采用一对多检索范式,聚焦未剪辑视频,适用于长视频平台场景。2)多层级视觉对应:为涵盖新闻、旅行、舞蹈等常见视频类别,基于用户关注的核心内容(如新闻事件、旅行地点、舞蹈动作)构建六级视觉对应体系,包括复制、事件、场景、实例、动作和其他层级。3)全面评估标准:设计三个版本的MUVR(Base、Filter、QA),其中MUVR-Base/Filter用于评估检索模型,而MUVR-QA以问答形式评估多模态大语言模型(MLLMs);同时提出重排序评分以评估MLLM的重排序能力。MUVR包含来自Bilibili平台的53,000段未剪辑视频,1,050个多模态查询和84,000个匹配关系。对3个先进视频检索模型、6个基于图像的VLMs和10个MLLMs进行了广泛评估。结果揭示了现有方法在处理未剪辑视频和多模态查询方面的局限性,以及MLLMs在多视频理解与重排序上的不足。代码与数据集已公开于https://github.com/debby-0527/MUVR。

原文摘要 · Abstract (English)

We propose the Multi-modal Untrimmed Video Retrieval task, along with a new benchmark (MUVR) to advance video retrieval for long-video platforms. MUVR aims to retrieve untrimmed videos containing relevant segments using multi-modal queries. It has the following features: 1) Practical retrieval paradigm: MUVR supports video-centric multi-modal queries, expressing fine-grained retrieval needs through long text descriptions, video tag prompts, and mask prompts. It adopts a one-to-many retrieval paradigm and focuses on untrimmed videos, tailored for long-video platform applications. 2) Multi-level visual correspondence: To cover common video categories (e.g., news, travel, dance) and precisely define retrieval matching criteria, we construct multi-level visual correspondence based on core video content (e.g., news events, travel locations, dance moves) which users are interested in and want to retrieve. It covers six levels: copy, event, scene, instance, action, and others. 3) Comprehensive evaluation criteria: We develop 3 versions of MUVR (i.e., Base, Filter, QA). MUVR-Base/Filter evaluates retrieval models, while MUVR-QA assesses MLLMs in a question-answering format. We also propose a Reranking Score to evaluate the reranking ability of MLLMs. MUVR consists of 53K untrimmed videos from the video platform Bilibili, with 1,050 multi-modal queries and 84K matches. Extensive evaluations of 3 state-of-the-art video retrieval models, 6 image-based VLMs, and 10 MLLMs are conducted. MUVR reveals the limitations of retrieval methods in processing untrimmed videos and multi-modal queries, as well as MLLMs in multi-video understanding and reranking. Our code and benchmark is available at https://github.com/debby-0527/MUVR.

视频检索多模态长视频基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。