arXiv:2512.10945cs.CV2025-12TPAMI被引 64

构建首个聚焦运动表达的视频分割数据集,推动基于语言描述的动态目标理解。

MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation

  • 构建包含3.3万条人标注运动表达的多模态数据集,覆盖2006段复杂场景视频
  • 15种现有方法在4项任务上表现不佳,暴露对运动语义理解的不足
  • 提出LMPM++新方法,在视频分割与追踪任务中达最新水平,适合动态视觉理解研究者

本文提出一个大规模多模态数据集MeViS,用于基于语言描述的运动表达视频分割,旨在根据物体运动的语言描述实现视频中目标对象的像素级分割与跟踪。现有引用视频分割数据集多关注显著对象,语言表达侧重静态属性,可能导致仅凭单帧即可识别目标,忽视了视频与语言中运动的重要性。为探索运动表达与运动推理线索在像素级视频理解中的可行性,我们构建了包含33,072条人工标注的文本与语音运动表达、涵盖8,171个物体的2,006段复杂场景视频的MeViS数据集。我们在该数据集上对15种现有方法进行了基准测试,覆盖4项任务:6种引用视频对象分割(RVOS)方法、3种音频引导视频对象分割(AVOS)方法、2种引用多目标追踪(RMOT)方法,以及4种视频描述生成方法用于新提出的引用运动表达生成(RMEG)任务。结果表明,现有方法在处理运动表达引导的视频理解方面存在明显弱点。我们进一步分析挑战并提出一种改进方法LMPM++,在RVOS/AVOS/RMOT任务上取得新最佳性能。本数据集为复杂视频场景中运动表达引导的视频理解算法发展提供了平台。数据集及代码已公开于https://henghuiding.com/MeViS/

原文摘要 · Abstract (English)

This paper proposes a large-scale multi-modal dataset for referring motion expression video segmentation, focusing on segmenting and tracking target objects in videos based on language description of objects' motions. Existing referring video segmentation datasets often focus on salient objects and use language expressions rich in static attributes, potentially allowing the target object to be identified in a single frame. Such datasets underemphasize the role of motion in both videos and languages. To explore the feasibility of using motion expressions and motion reasoning clues for pixel-level video understanding, we introduce MeViS, a dataset containing 33,072 human-annotated motion expressions in both text and audio, covering 8,171 objects in 2,006 videos of complex scenarios. We benchmark 15 existing methods across 4 tasks supported by MeViS, including 6 referring video object segmentation (RVOS) methods, 3 audio-guided video object segmentation (AVOS) methods, 2 referring multi-object tracking (RMOT) methods, and 4 video captioning methods for the newly introduced referring motion expression generation (RMEG) task. The results demonstrate weaknesses and limitations of existing methods in addressing motion expression-guided video understanding. We further analyze the challenges and propose an approach LMPM++ for RVOS/AVOS/RMOT that achieves new state-of-the-art results. Our dataset provides a platform that facilitates the development of motion expression-guided video understanding algorithms in complex video scenes. The proposed MeViS dataset and the method's source code are publicly available at https://henghuiding.com/MeViS/

视频分割多模态运动表达数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。