arXiv:2512.17909cs.CV2025-12被引 19

让视觉理解模型的特征适配图像生成,兼顾语义与细节重建。

Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing

  • 设计语义-像素联合重建目标,压缩特征到96通道
  • 在16x16下采样空间中实现最优重建与生成效果
  • 适合需要统一生成与编辑能力的研究者使用

现代潜在扩散模型通常在以像素重建优化的低维变分自编码器(VAE)潜在空间中运行。为统一视觉生成与理解,近期趋势是采用面向理解的表示编码器高维特征作为生成潜变量。然而我们实证发现该范式存在两大障碍:(1) 判别性特征空间缺乏紧凑正则化,导致扩散模型易产生离曼达特潜在变量,造成物体结构失真;(2) 编码器本身像素级重建能力弱,阻碍生成器学习精确的细粒度几何与纹理。本文提出系统性框架,将理解导向的编码器特征适配生成任务。引入语义-像素重建目标,使语义信息与细粒度细节压缩至高度紧凑的表示(96通道,16×16空间下采样)。该设计确保潜在空间兼具语义丰富性与先进图像重建性能,同时保持紧凑性以支持精准生成。基于此表示,构建统一的文本到图像生成与图像编辑模型。在多种特征空间对比中,本方法在重建、收敛速度及生成与编辑任务上均达最优,验证了表示编码器可被有效转化为鲁棒生成组件。

原文摘要 · Abstract (English)

Modern Latent Diffusion Models (LDMs) typically operate in low-level Variational Autoencoder (VAE) latent spaces that are primarily optimized for pixel-level reconstruction. To unify vision generation and understanding, a burgeoning trend is to adopt high-dimensional features from representation encoders as generative latents. However, we empirically identify two fundamental obstacles in this paradigm: (1) the discriminative feature space lacks compact regularization, making diffusion models prone to off-manifold latents that lead to inaccurate object structures; and (2) the encoder's inherently weak pixel-level reconstruction hinders the generator from learning accurate fine-grained geometry and texture. In this paper, we propose a systematic framework to adapt understanding-oriented encoder features for generative tasks. We introduce a semantic-pixel reconstruction objective to regularize the latent space, enabling the compression of both semantic information and fine-grained details into a highly compact representation (96 channels with 16x16 spatial downsampling). This design ensures that the latent space remains semantically rich and achieves state-of-the-art image reconstruction, while remaining compact enough for accurate generation. Leveraging this representation, we design a unified Text-to-Image (T2I) and image editing model. Benchmarking against various feature spaces, we demonstrate that our approach achieves state-of-the-art reconstruction, faster convergence, and substantial performance gains in both T2I and editing tasks, validating that representation encoders can be effectively adapted into robust generative components.

图像生成特征对齐扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。