arXiv:2504.05468cs.CV2025-04CVPR被引 4

用扩散模型特征实现零样本视频目标分割,无需微调

Studying Image Diffusion Features for Zero-Shot Video Object Segmentation

  • 从扩散模型不同层和时间步提取特征,找到最优组合
  • 在DAVIS-17和MOSE上达到当前最好效果,性能媲美有监督模型
  • 点对应关系是高精度分割的关键,适合零样本场景

本文研究大规模扩散模型在零样本视频目标分割(ZS-VOS)中的应用,不需在视频数据上微调或使用任何图像分割训练数据。尽管扩散模型在各类任务中展现强大视觉表征能力,其在ZS-VOS中的直接应用仍鲜有探索。我们通过实验寻找最适合的特征提取策略,即选择最佳的时间步与网络层。进一步分析发现,提取的特征与点对应关系具有强相关性。在DAVIS-17和MOSE数据集上的大量实验证明:在ImageNet上训练的扩散模型,反而优于在更大更多样数据集上训练的模型。此外,点对应关系对高精度分割至关重要,本方法实现了当前最优的ZS-VOS性能,且表现可媲美依赖昂贵图像分割数据集训练的模型。

原文摘要 · Abstract (English)

This paper investigates the use of large-scale diffusion models for Zero-Shot Video Object Segmentation (ZS-VOS) without fine-tuning on video data or training on any image segmentation data. While diffusion models have demonstrated strong visual representations across various tasks, their direct application to ZS-VOS remains underexplored. Our goal is to find the optimal feature extraction process for ZS-VOS by identifying the most suitable time step and layer from which to extract features. We further analyze the affinity of these features and observe a strong correlation with point correspondences. Through extensive experiments on DAVIS-17 and MOSE, we find that diffusion models trained on ImageNet outperform those trained on larger, more diverse datasets for ZS-VOS. Additionally, we highlight the importance of point correspondences in achieving high segmentation accuracy, and we yield state-of-the-art results in ZS-VOS. Finally, our approach performs on par with models trained on expensive image segmentation datasets.

视频分割扩散模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。