arXiv:2604.05730cs.LG2026-04中稿 · CVPR

让图像生成精准组合多种输入条件,速度更快且可控性更强。

Controllable Image Generation with Composed Parallel Token Prediction

  • 通过可组合的并行令牌预测,实现多条件精准控制
  • 在3个数据集上误差降低63.4%,FID提升9.58
  • 支持细粒度文本生成控制,推理速度快2.3到12倍

条件离散生成模型难以忠实组合多个输入条件。为此,我们推导出一种理论完备的离散概率生成过程组合方法,其中掩码生成(吸收扩散)为特例。该方法能精确指定训练数据之外的新组合与条件数量,概念加权可强调或否定特定条件。结合VQ-VAE和VQ-GAN丰富的组合式语义词典,本方法在位置型CLEVR、关系型CLEVR和FFHQ三个数据集上,相对之前最优方法实现63.4%的误差率下降,平均绝对FID提升-9.58。同时,相比同类方法,推理速度提升2.3至12倍,并可直接应用于开源预训练离散文生图模型,实现细粒度控制。

原文摘要 · Abstract (English)

Conditional discrete generative models struggle to faithfully compose multiple input conditions. To address this, we derive a theoretically-grounded formulation for composing discrete probabilistic generative processes, with masked generation (absorbing diffusion) as a special case. Our formulation enables precise specification of novel combinations and numbers of input conditions that lie outside the training data, with concept weighting enabling emphasis or negation of individual conditions. In synergy with the richly compositional learned vocabulary of VQ-VAE and VQ-GAN, our method attains a $63.4\%$ relative reduction in error rate compared to the previous state-of-the-art, averaged across 3 datasets (positional CLEVR, relational CLEVR and FFHQ), simultaneously obtaining an average absolute FID improvement of $-9.58$. Meanwhile, our method offers a $2.3\times$ to $12\times$ real-time speed-up over comparable methods, and is readily applied to an open pre-trained discrete text-to-image model for fine-grained control of text-to-image generation.

图像生成条件控制离散生成加速推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。