提出自控机制,让图像生成更精准可控。
Self-control: A Better Conditional Mechanism for Masked Autoregressive Model
- 用自注意力机制融合文本图像条件,替代传统交叉注意力
- 在连续掩码自回归模型中实现多模态条件统一建模
- 避免向量量化干扰,提升生成图像质量与条件控制力
自回归条件图像生成算法能够生成与给定文本或图像条件一致的逼真图像,在多种应用中具有巨大潜力。然而,多数主流方法严重依赖向量量化,其固有的离散特性对高质量图像生成构成显著挑战。为此,本文提出一种针对连续掩码自回归模型的新式条件引入网络——自控网络。该网络旨在缓解向量量化对生成图像质量的负面影响,同时增强生成过程中的条件控制能力。具体而言,自控网络基于连续掩码自回归生成模型,以串行方式将文本和图像等多模态条件信息融入统一的自回归序列中。通过自注意力机制,模型可依据特定条件生成可控图像。该网络摒弃传统的基于交叉注意力的条件融合方式,有效将条件信息与生成信息统一于同一空间,从而促进多模态特征更无缝地学习与融合。
原文摘要 · Abstract (English)
Autoregressive conditional image generation algorithms are capable of generating photorealistic images that are consistent with given textual or image conditions, and have great potential for a wide range of applications. Nevertheless, the majority of popular autoregressive image generation methods rely heavily on vector quantization, and the inherent discrete characteristic of codebook presents a considerable challenge to achieving high-quality image generation. To address this limitation, this paper introduces a novel conditional introduction network for continuous masked autoregressive models. The proposed self-control network serves to mitigate the negative impact of vector quantization on the quality of the generated images, while simultaneously enhancing the conditional control during the generation process. In particular, the self-control network is constructed upon a continuous mask autoregressive generative model, which incorporates multimodal conditional information, including text and images, into a unified autoregressive sequence in a serial manner. Through a self-attention mechanism, the network is capable of generating images that are controllable based on specific conditions. The self-control network discards the conventional cross-attention-based conditional fusion mechanism and effectively unifies the conditional and generative information within the same space, thereby facilitating more seamless learning and fusion of multimodal features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。