arXiv:2510.06139cs.CV2025-10被引 8

用连续形变建模视频分割,让语言描述直接引导像素变化。

Deforming Videos to Masks: Flow Matching for Referring Video Segmentation

  • 将语言引导的视频分割视为从视频整体到目标掩码的连续形变过程。
  • 在MeViS上达到51.1的J&F,零样本下Ref-DAVIS17达73.3,刷新纪录。
  • 无需分步定位与分割,适合追求高精度与时序一致性的研究者。

指代式视频目标分割(RVOS)需根据自然语言描述分割视频中的特定对象。核心挑战在于将抽象的语言概念锚定到具体像素,并在复杂视频动态中持续跟踪。以往方法常采用‘定位-分割’的级联流程,但这种设计将语义简化为粗粒度几何提示(如点),导致信息瓶颈,且难以保持时序一致性。为此,本文提出FlowRVS,将RVOS重构为条件连续流问题,充分利用预训练文本到视频模型的优势、精细像素控制、文本-视频语义对齐及时间连贯性。不同于传统从噪声生成掩码或直接预测掩码,本方法学习从视频整体表示到目标掩码的直接、语言引导的形变过程。该单阶段生成式框架在所有主要RVOS基准上均达到新最优性能:在MeViS上取得51.1的J&F(较之前SOTA提升1.6),在零样本设置下的Ref-DAVIS17上达到73.3(提升2.7),验证了将视频理解任务建模为连续形变过程的巨大潜力。

原文摘要 · Abstract (English)

Referring Video Object Segmentation (RVOS) requires segmenting specific objects in a video guided by a natural language description. The core challenge of RVOS is to anchor abstract linguistic concepts onto a specific set of pixels and continuously segment them through the complex dynamics of a video. Faced with this difficulty, prior work has often decomposed the task into a pragmatic `locate-then-segment' pipeline. However, this cascaded design creates an information bottleneck by simplifying semantics into coarse geometric prompts (e.g, point), and struggles to maintain temporal consistency as the segmenting process is often decoupled from the initial language grounding. To overcome these fundamental limitations, we propose FlowRVS, a novel framework that reconceptualizes RVOS as a conditional continuous flow problem. This allows us to harness the inherent strengths of pretrained T2V models, fine-grained pixel control, text-video semantic alignment, and temporal coherence. Instead of conventional generating from noise to mask or directly predicting mask, we reformulate the task by learning a direct, language-guided deformation from a video's holistic representation to its target mask. Our one-stage, generative approach achieves new state-of-the-art results across all major RVOS benchmarks. Specifically, achieving a J&F of 51.1 in MeViS (+1.6 over prior SOTA) and 73.3 in the zero shot Ref-DAVIS17 (+2.7), demonstrating the significant potential of modeling video understanding tasks as continuous deformation processes.

视频分割语言引导连续形变多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。