arXiv:2605.27343cs.CVcs.LG2026-05

用自监督模型表示条件控制图像生成,提升可控性与质量

Towards Controllable Image Generation through Representation-Conditioned Diffusion Models

论文配图:Towards Controllable Image Generation through Representation-Conditioned Diffusion Models
图 1 · 摘自论文原文
  • 用预训练自监督模型的表征作为扩散模型的条件
  • 生成图像质量提升,且表征空间具平滑与解耦特性
  • 适合研究可控生成与无标注数据条件下的图像合成

扩散模型已成为高质量图像生成与编辑的强大工具,但引导其生成特定输出仍具挑战。传统方法依赖文本提示或语义图等条件机制,需大量标注数据。本文初步探索了基于预训练自监督模型表征的扩散模型条件化方法。自条件机制不仅提升了无条件图像生成的质量,还构建了一个可用于控制生成的表征空间。通过识别变化方向,实验表明该空间具有良好的平滑性与解耦性,展现出可控生成的潜力。

原文摘要 · Abstract (English)

Diffusion models have emerged as powerful tools for high-quality image generation and editing, but guiding these models to produce specific outputs remains a challenge. Conventional approaches rely on conditioning mechanisms, such as text prompts or semantic maps, which require extensively annotated datasets. In this preliminary work, we explore diffusion models conditioned on representations from a pre-trained self-supervised model. The self-conditioning mechanism not only improves the quality of unconditional image generation, but also provides a representation space that can be used to control the generation. We explore this conditioning space by identifying directions of variations, and demonstrate promising properties in terms of smoothness and disentanglement.

扩散模型可控生成自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。