一歩で高精細な画像復元を実現する軽量モデルGenDRの提案
GenDR: Lighten Generative Detail Restoration
- 1ステップの拡散モデルで詳細復元を実現、推論速度と精度の両立を達成
- 0.9Bパラメータの新規16チャネルVAEとスコア蒸留で、高周波成分を正確に再構成
- 生成品質向上と学習加速を両立する対抗的蒸留手法、実用性に優れる
尽管基于文本到图像(T2I)扩散模型的现实世界超分辨率(SR)研究取得了显著进展,但其目标不匹配导致推理速度与细节保真度之间存在次优权衡。具体而言,T2I任务需多步推断以匹配提示,并降低潜在维度以减轻生成难度;而SR可在较少步骤内恢复高频细节,但需要更可靠的变分自编码器(VAE)以保留输入信息。然而,多数基于扩散的SR为多步且使用4通道VAE,而现有16通道VAE模型如FLUX(12B)过于庞大。为此,我们提出从定制扩散模型中蒸馏出的一步式生成细节恢复模型GenDR,该模型具有更大的潜在空间。具体地,我们通过表示对齐训练新的SD2.1-VAE16(0.9B),在不增加模型规模的情况下扩展潜在空间。关于步骤蒸馏,我们提出一致得分身份蒸馏(CiD),将针对SR任务的损失融入得分蒸馏,以利用更多SR先验并对齐训练目标。此外,我们通过引入对抗学习和表示对齐扩展了CiD(CiDA),以提升感知质量并加速训练。同时优化流水线以实现更高效的推理。实验结果表明,GenDR在定量指标和视觉保真度方面均达到最先进水平。
原文摘要 · Abstract (English)
Although recent research applying text-to-image (T2I) diffusion models to real-world super-resolution (SR) has achieved remarkable progress, the misalignment of their targets leads to a suboptimal trade-off between inference speed and detail fidelity. Specifically, the T2I task requires multiple inference steps to synthesize images matching to prompts and reduces the latent dimension to lower generating difficulty. Contrariwise, SR can restore high-frequency details in fewer inference steps, but it necessitates a more reliable variational auto-encoder (VAE) to preserve input information. However, most diffusion-based SRs are multistep and use 4-channel VAEs, while existing models with 16-channel VAEs are overqualified diffusion transformers, e.g., FLUX (12B). To align the target, we present a one-step diffusion model for generative detail restoration, GenDR, distilled from a tailored diffusion model with a larger latent space. In detail, we train a new SD2.1-VAE16 (0.9B) via representation alignment to expand the latent space without increasing the model size. Regarding step distillation, we propose consistent score identity distillation (CiD) that incorporates SR task-specific loss into score distillation to leverage more SR priors and align the training target. Furthermore, we extend CiD with adversarial learning and representation alignment (CiDA) to enhance perceptual quality and accelerate training. We also polish the pipeline to achieve a more efficient inference. Experimental results demonstrate that GenDR achieves state-of-the-art performance in both quantitative metrics and visual fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。