用噪声提升蒙版生成图像的多样性与形态一致性
Diffusion Prism: Enhancing Diversity and Morphology Consistency in Mask-to-Image Diffusion
- 引入微量人工噪声增强去噪过程,提升生成多样性
- 在纳米树状结构上生成多样且形态一致的图像,优于现有模型
- 无需训练,适用于生物图案等复杂形态生成
生成式AI与可控扩散模型的发展使图像到图像的合成日益高效实用。然而,当输入图像熵值低且信息稀疏时,扩散模型固有的特性常导致生成多样性受限,严重影响数据增强效果。为此,我们提出Diffusion Prism,一个无需训练的框架,可将二值蒙版高效转换为真实且多样的样本,同时保持形态特征。我们发现少量人工噪声能显著辅助图像去噪过程。通过纳米树状结构作为示例,验证了该方法相比现有可控扩散模型的优势。此外,我们将该框架扩展至其他生物图案,展示了其在多个领域的应用潜力。
原文摘要 · Abstract (English)
The emergence of generative AI and controllable diffusion has made image-to-image synthesis increasingly practical and efficient. However, when input images exhibit low entropy and sparse, the inherent characteristics of diffusion models often result in limited diversity. This constraint significantly interferes with data augmentation. To address this, we propose Diffusion Prism, a training-free framework that efficiently transforms binary masks into realistic and diverse samples while preserving morphological features. We explored that a small amount of artificial noise will significantly assist the image-denoising process. To prove this novel mask-to-image concept, we use nano-dendritic patterns as an example to demonstrate the merit of our method compared to existing controllable diffusion models. Furthermore, we extend the proposed framework to other biological patterns, highlighting its potential applications across various fields.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。