arXiv:2604.17492cs.CV2026-04

让语义表示随生成任务动态演化,提升图像扩散模型质量

Coevolving Representations in Joint Image-Feature Diffusion

论文配图:Coevolving Representations in Joint Image-Feature Diffusion
图 1 · 摘自论文原文
  • 训练时同步优化语义表示与扩散模型,实现表示空间自适应
  • 相比固定表示,收敛更快且生成图像质量更高
  • 适合追求高质量图像生成与表示学习融合的研究者

联合图像-特征生成建模近年成为提升扩散训练的有效策略,通过将低层VAE隐变量与预训练视觉编码器提取的高层语义特征耦合。然而现有方法依赖独立构建且固定不变的表示空间。本文提出共演化表示扩散(CoReDi),在训练中通过轻量线性投影联合学习语义表示空间。为防止退化,引入停止梯度目标、归一化和针对性正则化以维持稳定性。该框架使语义空间逐步适配生成需求,增强其与图像隐变量的互补性。应用于VAE隐空间与像素空间扩散模型均取得改进,实验表明相比固定表示,收敛更快且样本质量更高。

原文摘要 · Abstract (English)

Joint image-feature generative modeling has recently emerged as an effective strategy for improving diffusion training by coupling low-level VAE latents with high-level semantic features extracted from pre-trained visual encoders. However, existing approaches rely on a fixed representation space, constructed independently of the generative objective and kept unchanged during training. We argue that the representation space guiding diffusion should itself adapt to the generative task. To this end, we propose Coevolving Representation Diffusion (CoReDi), a framework in which the semantic representation space evolves during training by learning a lightweight linear projection jointly with the diffusion model. While naively optimizing this projection leads to degenerate solutions, we show that stable coevolution can be achieved through a combination of stop-gradient targets, normalization, and targeted regularization that prevents feature collapse. This formulation enables the semantic space to progressively specialize to the needs of image synthesis, improving its complementarity with image latents. We apply CoReDi to both VAE latent diffusion and pixel-space diffusion, demonstrating that adaptive semantic representations improve generative modeling across both settings. Experiments show that CoReDi achieves faster convergence and higher sample quality compared to joint diffusion models operating in fixed representation spaces.

扩散模型表示学习图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。