arXiv:2508.13584cs.CV2025-08被引 1

用无噪扩散模型提升视频目标分割精度,边界更准。

Temporal-Conditional Referring Video Object Segmentation with Noise-Free Text-to-Video Diffusion Model

  • 用无噪扩散模型提取特征,避免噪声干扰
  • 设计轻量级时序上下文掩码精修模块,边界分割更优
  • 在4个公开数据集上达到顶尖性能,适合视频理解任务

指代式视频目标分割(RVOS)旨在根据文本描述定位视频中的特定对象。我们发现现有方法过度关注特征提取与时序建模,而忽视了分割头的设计。为此,提出一种时序条件型指代式视频目标分割模型,创新性地融合现有分割方法以增强边界分割能力。同时,采用文本到视频扩散模型进行特征提取,并移除传统噪声预测模块,避免噪声引入随机性对分割精度的负面影响,简化模型结构的同时提升性能。此外,针对VAE特征提取能力有限的问题,设计时序上下文掩码精修(TCMR)模块,在不增加复杂度的前提下显著提升分割质量。在四个公开的RVOS基准上评估,该方法持续取得领先表现。

原文摘要 · Abstract (English)

Referring Video Object Segmentation (RVOS) aims to segment specific objects in a video according to textual descriptions. We observe that recent RVOS approaches often place excessive emphasis on feature extraction and temporal modeling, while relatively neglecting the design of the segmentation head. In fact, there remains considerable room for improvement in segmentation head design. To address this, we propose a Temporal-Conditional Referring Video Object Segmentation model, which innovatively integrates existing segmentation methods to effectively enhance boundary segmentation capability. Furthermore, our model leverages a text-to-video diffusion model for feature extraction. On top of this, we remove the traditional noise prediction module to avoid the randomness of noise from degrading segmentation accuracy, thereby simplifying the model while improving performance. Finally, to overcome the limited feature extraction capability of the VAE, we design a Temporal Context Mask Refinement (TCMR) module, which significantly improves segmentation quality without introducing complex designs. We evaluate our method on four public RVOS benchmarks, where it consistently achieves state-of-the-art performance.

视频分割扩散模型文本-视频对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。