arXiv:2508.04122cs.CV2025-08ICCV被引 5

用扩散模型实现零样本实例分割,无需训练即能精准识别物体

Conditional Latent Diffusion Models for Zero-Shot Instance Segmentation

  • 在隐空间中通过物体模板和图像特征引导生成实例掩码
  • 在多个真实世界数据集上达到当前最佳性能,且无需目标数据微调
  • 适合研究零样本学习与生成式分割的学者参考

本文提出OC-DiT,一种面向以物体为中心预测的新式扩散模型,并将其应用于零样本实例分割。我们设计了一种条件隐空间扩散框架,通过在扩散过程的隐空间中结合物体模板与图像特征来生成实例掩码,从而有效利用视觉物体描述符和局部图像线索分离不同物体实例。具体地,提出两种模型变体:一个粗略模型用于生成初始实例提议,一个精修模型并行优化所有提议。模型在包含数千个高质量物体网格的全新大规模合成数据集上训练。令人瞩目的是,该模型在多个挑战性真实世界基准上实现了当前最优表现,且无需在目标数据上进行任何微调。通过全面的消融实验,验证了扩散模型在实例分割任务中的潜力。

原文摘要 · Abstract (English)

This paper presents OC-DiT, a novel class of diffusion models designed for object-centric prediction, and applies it to zero-shot instance segmentation. We propose a conditional latent diffusion framework that generates instance masks by conditioning the generative process on object templates and image features within the diffusion model's latent space. This allows our model to effectively disentangle object instances through the diffusion process, which is guided by visual object descriptors and localized image cues. Specifically, we introduce two model variants: a coarse model for generating initial object instance proposals, and a refinement model that refines all proposals in parallel. We train these models on a newly created, large-scale synthetic dataset comprising thousands of high-quality object meshes. Remarkably, our model achieves state-of-the-art performance on multiple challenging real-world benchmarks, without requiring any retraining on target data. Through comprehensive ablation studies, we demonstrate the potential of diffusion models for instance segmentation tasks.

实例分割扩散模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。