通过自引导提升掩码生成模型图像质量,兼顾多样性与效率
Unlocking the Capabilities of Masked Generative Models for Image Synthesis via Self-Guidance
- 提出自引导采样方法,利用语义平滑增强生成质量
- 在相同参数量下超越现有掩码生成模型,质量与多样性更优
- 训练和采样成本更低,适合高效图像生成场景
掩码生成模型(MGMs)在生成能力上表现优异,采样步数比连续扩散模型低一个数量级。然而,在生成质量与多样性方面,仍逊于同规模的先进连续扩散模型。连续扩散模型性能优势主要来自引导方法,但会牺牲多样性。本文将此类引导方法推广至MGMs的通用形式,提出自引导采样策略:在向量量化标记空间中引入辅助任务实现语义平滑,类比于连续像素空间中的高斯模糊。结合参数高效微调与高温采样,所提方法在保持较低训练与采样成本的同时,显著提升生成质量与多样性平衡。大量实验验证了该方法在不同采样超参数下的有效性。
原文摘要 · Abstract (English)
Masked generative models (MGMs) have shown impressive generative ability while providing an order of magnitude efficient sampling steps compared to continuous diffusion models. However, MGMs still underperform in image synthesis compared to recent well-developed continuous diffusion models with similar size in terms of quality and diversity of generated samples. A key factor in the performance of continuous diffusion models stems from the guidance methods, which enhance the sample quality at the expense of diversity. In this paper, we extend these guidance methods to generalized guidance formulation for MGMs and propose a self-guidance sampling method, which leads to better generation quality. The proposed approach leverages an auxiliary task for semantic smoothing in vector-quantized token space, analogous to the Gaussian blur in continuous pixel space. Equipped with the parameter-efficient fine-tuning method and high-temperature sampling, MGMs with the proposed self-guidance achieve a superior quality-diversity trade-off, outperforming existing sampling methods in MGMs with more efficient training and sampling costs. Extensive experiments with the various sampling hyperparameters confirm the effectiveness of the proposed self-guidance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。