arXiv:2411.01819cs.CVcs.AI2024-11中稿 · ACM MM被引 6

用扩散模型自动生成多实例分割数据,提升真实感与标注精度。

Free-Mask: A Novel Paradigm of Integration Between the Segmentation Diffusion Model and Image Editing

  • 将分割扩散模型与图像编辑结合,通过文本生成多物体图像。
  • 合成数据训练的模型在VOC 2012上对未见类别实现新SOTA。
  • 适合需要低成本高质量分割数据集的研究者使用。

当前语义分割模型通常依赖大量人工标注数据,耗时且成本高。利用Midjourney和Stable Diffusion等先进文本到图像模型生成合成数据成为高效替代方案。然而,以往方法仅能生成单实例图像,因Stable Diffusion生成多实例不稳定而受限。为此,我们提出全新框架Free-Mask,将分割扩散模型与高级图像编辑能力融合,通过文本到图像模型实现多物体图像的集成。该方法可生成高度逼真的数据集,贴近开放世界环境,并自动生成精确分割掩码。显著降低人工标注工作量,同时保证掩码准确性。实验表明,由Free-Mask生成的合成数据使分割模型在零样本设置下表现超越真实数据训练模型。尤其在VOC 2012基准上,对先前未见类别的性能达到新SOTA。

原文摘要 · Abstract (English)

Current semantic segmentation models typically require a substantial amount of manually annotated data, a process that is both time-consuming and resource-intensive. Alternatively, leveraging advanced text-to-image models such as Midjourney and Stable Diffusion has emerged as an efficient strategy, enabling the automatic generation of synthetic data in place of manual annotations. However, previous methods have been limited to generating single-instance images, as the generation of multiple instances with Stable Diffusion has proven unstable. To address this limitation and expand the scope and diversity of synthetic datasets, we propose a framework \textbf{Free-Mask} that combines a Diffusion Model for segmentation with advanced image editing capabilities, allowing for the integration of multiple objects into images via text-to-image models. Our method facilitates the creation of highly realistic datasets that closely emulate open-world environments while generating accurate segmentation masks. It reduces the labor associated with manual annotation and also ensures precise mask generation. Experimental results demonstrate that synthetic data generated by \textbf{Free-Mask} enables segmentation models to outperform those trained on real data, especially in zero-shot settings. Notably, \textbf{Free-Mask} achieves new state-of-the-art results on previously unseen classes in the VOC 2012 benchmark.

分割生成扩散模型合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。