arXiv:2606.05652cs.CV2026-06

无需标签即可分步生成图像,控制更精准。

CoFi-UCGen: Coarse-to-Fine Unsupervised Conditional Generation without Label Priors

论文配图:CoFi-UCGen: Coarse-to-Fine Unsupervised Conditional Generation without Label Priors
图 1 · 摘自论文原文
  • 先学粗粒度语义再逐步细化,显式分离全局与细节特征。
  • 在无标签情况下生成质量、一致性、可控性均超越现有方法。
  • 适合需要精准图像控制但无标注数据的生成任务。

无监督条件图像生成(UCGen)旨在不依赖人工标注标签的情况下实现生成控制,但因跨粒度语义表示不结构化而面临挑战。为此,我们提出首个无需任何标签先验的粗到细无监督条件生成框架(CoFi-UCGen),显式解耦全局语义与细粒度变化。我们首先提出对抗性语义互学习理论,确保图像与潜在空间间语义一致性和完整性;基于此,引入位码(bit-codes)构建结构化的粗粒度潜在空间,并证明其蕴含独立的全局语义,同时保留独立噪声采样以支持生成。在此基础上,建立细粒度语义基,设计扩散模型中的分层调制机制,通过逐层注入粗粒度条件,逐步控制生成过程中的细粒度属性。大量实验表明,在无标签先验或预训练特征提取器条件下,我们的CoFi-UCGen在图像质量、语义一致性与控制准确性上持续优于现有方法,验证了显式粗到细语义分解在挑战性无监督条件生成任务中的有效性。

原文摘要 · Abstract (English)

Unsupervised conditional image generation (UCGen) aims to control generation without relying on manually annotated labels, yet remains challenging due to unstructured semantic representations across granularities. To address this, we propose a novel coarse-to-fine UCGen framework (CoFi-UCGen) that explicitly disentangles global semantics from fine-grained variations, which to the best of our knowledge, sets out the first successful attempt for both coarse- and fine-grained conditional generation without any labels. More specifically, we first propose the adversarial semantic reciprocal learning theory to ensure the semantic consistency and completeness between images and latent spaces. Based on the consistency, we propose the bit-codes to learn a structured coarse-grained latent space, and further prove distinct global semantics inherent from our bit-codes while preserving independent noise sampling for generation. Building upon these bit-codes, we establish a fine-grained semantic basis and introduce a hierarchical modulation mechanism in diffusion models, by enabling layer-wise injection from coarse conditions to progressively control fine-grained attributes during generation. Extensive experiments demonstrate that without any label priors or pre-trained feature extractors, our CoFi-UCGen consistently outperforms existing UCGen methods in terms of image quality, semantic consistency, and control accuracy, verifying the effectiveness of explicit coarse-to-fine semantic decomposition for the challenging UCGen task.

无监督生成条件生成扩散模型粗细粒度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。