arXiv:2602.22150cs.CV2026-02中稿 · CVPR被引 3

提出渐进式框架,统一图像生成中概念与定位的矛盾需求。

CoLoGen: Progressive Learning of Concept-Localization Duality for Unified Image Generation

  • 分阶段训练:先建模概念与定位能力,再适配条件,最后融合协同。
  • 核心模块动态路由特征,稳定整合多专家输出,避免表征冲突。
  • 在编辑、可控生成等任务上表现优越,适合复杂指令生成场景。

统一的条件图像生成仍具挑战,因不同任务依赖根本不同的内部表征:部分需语义理解以实现概念合成,另一些则依赖空间定位以保证精度。强制这些异构任务共享单一表征会引发概念-定位表征冲突。为此,我们提出CoLoGen,一种统一的扩散框架,可渐进学习并调和这种概念-定位二元性。CoLoGen采用分阶段课程策略:首先构建核心概念与定位能力,然后将其适应到多样视觉条件,最后精炼其协同关系以应对复杂指令驱动任务。该过程的核心是渐进表征编织(PRW)模块,能动态路由特征至专用专家,并在各阶段稳定整合其输出。在图像编辑、可控生成和定制化生成任务上的实验表明,CoLoGen性能达到或优于现有方法,为统一图像生成提供了原理性的表征视角。

原文摘要 · Abstract (English)

Unified conditional image generation remains difficult because different tasks depend on fundamentally different internal representations. Some require conceptual understanding for semantic synthesis, while others rely on localization cues for spatial precision. Forcing these heterogeneous tasks to share a single representation leads to concept-localization representational conflict. To address this issue, we propose CoLoGen, a unified diffusion framework that progressively learns and reconciles this concept-localization duality. CoLoGen uses a staged curriculum that first builds core conceptual and localization abilities, then adapts them to diverse visual conditions, and finally refines their synergy for complex instruction-driven tasks. Central to this process is the Progressive Representation Weaving (PRW) module, which dynamically routes features to specialized experts and stably integrates their outputs across stages. Experiments on editing, controllable generation, and customized generation show that CoLoGen achieves competitive or superior performance, offering a principled representational perspective for unified image generation.

图像生成扩散模型统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。