arXiv:2412.19080cs.CV2024-12NeurIPS被引 17

用新方法生成高质量图像分割数据,省时省钱还精准。

Mask Factory: Towards High-quality Synthetic Data Generation for Dichotomous Image Segmentation

  • 结合刚性与非刚性编辑生成高精度合成掩码
  • 在DIS5K数据集上效果优于现有方法,效率显著提升
  • 适合需要大量标注数据的视觉研究者使用

二值图像分割(DIS)任务需要高度精确的标注,传统数据集构建方法耗时、成本高且依赖领域知识。尽管合成数据是潜在解决方案,但现有生成模型常面临场景偏差、噪声干扰及训练样本多样性不足的问题。为此,我们提出一种新方法 extbf{ heirmodel{}},可高效生成多样且精确的数据集,显著降低准备时间和成本。首先,提出一种通用掩码编辑方法,融合刚性与非刚性编辑:刚性编辑利用扩散模型的几何先验,在零样本条件下实现精准视角变换;非刚性编辑则通过对抗训练与自注意力机制完成复杂且拓扑一致的修改。随后,采用多条件控制生成方法,输出高分辨率图像与准确分割掩码对。在广泛使用的DIS5K数据集基准测试中,本方法在质量和效率方面均优于现有技术。代码已开源:https://qian-hao-tian.github.io/MaskFactory/

原文摘要 · Abstract (English)

Dichotomous Image Segmentation (DIS) tasks require highly precise annotations, and traditional dataset creation methods are labor intensive, costly, and require extensive domain expertise. Although using synthetic data for DIS is a promising solution to these challenges, current generative models and techniques struggle with the issues of scene deviations, noise-induced errors, and limited training sample variability. To address these issues, we introduce a novel approach, \textbf{\ourmodel{}}, which provides a scalable solution for generating diverse and precise datasets, markedly reducing preparation time and costs. We first introduce a general mask editing method that combines rigid and non-rigid editing techniques to generate high-quality synthetic masks. Specially, rigid editing leverages geometric priors from diffusion models to achieve precise viewpoint transformations under zero-shot conditions, while non-rigid editing employs adversarial training and self-attention mechanisms for complex, topologically consistent modifications. Then, we generate pairs of high-resolution image and accurate segmentation mask using a multi-conditional control generation method. Finally, our experiments on the widely-used DIS5K dataset benchmark demonstrate superior performance in quality and efficiency compared to existing methods. The code is available at \url{https://qian-hao-tian.github.io/MaskFactory/}.

图像分割合成数据扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。