改进扩散模型生成质量与效率,支持连续标签可控图像生成。
Enhancing Diffusion-Based Quantitatively Controllable Image Generation via Matrix-Form EDM and Adaptive Vicinal Training
- 采用矩阵形式的EDM框架和自适应邻域训练策略
- 在4个数据集上实现更高图像质量,采样成本显著降低
- 适合需要高精度可控图像生成的研究者与开发者
连续条件扩散模型(CCDM)是一种基于扩散框架、可依据连续回归标签生成高质量图像的方法。尽管其在多个数据集上表现优于先前方法,但仍存在依赖过时扩散框架及采样轨迹过长导致效率低下的问题,近期更被基于GAN的CcGAN-AVAR方法超越。为此,本文提出改进版CCDM框架iCCDM,融合更先进的阐明扩散模型(EDM)框架,并引入新型矩阵形式的EDM表述与自适应邻域训练策略。在四个基准数据集(图像分辨率从64×64到256×256)上的大量实验表明,iCCDM持续优于现有方法,包括当前领先的大规模文本到图像扩散模型(如Stable Diffusion 3、FLUX.1和Qwen-Image),在提升生成质量的同时显著降低采样开销。
原文摘要 · Abstract (English)
Continuous Conditional Diffusion Model (CCDM) is a diffusion-based framework designed to generate high-quality images conditioned on continuous regression labels. Although CCDM has demonstrated clear advantages over prior approaches across a range of datasets, it still exhibits notable limitations and has recently been surpassed by a GAN-based method, namely CcGAN-AVAR. These limitations mainly arise from its reliance on an outdated diffusion framework and its low sampling efficiency due to long sampling trajectories. To address these issues, we propose an improved CCDM framework, termed iCCDM, which incorporates the more advanced \textit{Elucidated Diffusion Model} (EDM) framework with substantial modifications to improve both generation quality and sampling efficiency. Specifically, iCCDM introduces a novel matrix-form EDM formulation together with an adaptive vicinal training strategy. Extensive experiments on four benchmark datasets, spanning image resolutions from $64\times64$ to $256\times256$, demonstrate that iCCDM consistently outperforms existing methods, including state-of-the-art large-scale text-to-image diffusion models (e.g., Stable Diffusion 3, FLUX.1, and Qwen-Image), achieving higher generation quality while significantly reducing sampling cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。