arXiv:2507.23620cs.CVcs.LG2025-07AAAI被引 9

让图像生成模型轻松应对新控制条件,训练成本降36倍

DivControl: Knowledge Diversion for Controllable Image Generation

  • 通过分解ControlNet并分离通用与特定组件,实现可插拔式控制
  • 训练成本降低36.4倍,零样本泛化能力显著提升
  • 适合需要快速适配新控制条件的研究与应用

扩散模型已从文本到图像(T2I)发展为结合深度图等结构化输入的图像到图像(I2I)生成,实现精细空间控制。然而,现有方法要么为每种条件训练独立模型,要么依赖统一架构中的纠缠表征,导致泛化能力差且适应新条件成本高。为此,我们提出DivControl,一种可分解的预训练框架,支持统一可控生成与高效适应。DivControl通过奇异值分解(SVD)将ControlNet分解为基本分量——奇异向量对,并在多条件训练中通过知识分流机制,将其解耦为与条件无关的learnge(通用基因)和条件相关的tailor(专用配件)。知识分流由动态门控实现,根据条件指令语义对tailor进行软路由,从而实现零样本泛化和参数高效的新型条件适应。为进一步提升条件保真度与训练效率,引入表示对齐损失,使条件嵌入与扩散模型早期特征对齐。大量实验表明,DivControl在保持基础条件平均性能的同时,实现了最先进的可控性,训练成本降低36.4倍;对未见条件也展现出强大的零样本与少样本性能,验证了其卓越的可扩展性、模块化与迁移能力。

原文摘要 · Abstract (English)

Diffusion models have advanced from text-to-image (T2I) to image-to-image (I2I) generation by incorporating structured inputs such as depth maps, enabling fine-grained spatial control. However, existing methods either train separate models for each condition or rely on unified architectures with entangled representations, resulting in poor generalization and high adaptation costs for novel conditions. To this end, we propose DivControl, a decomposable pretraining framework for unified controllable generation and efficient adaptation. DivControl factorizes ControlNet via SVD into basic components-pairs of singular vectors-which are disentangled into condition-agnostic learngenes and condition-specific tailors through knowledge diversion during multi-condition training. Knowledge diversion is implemented via a dynamic gate that performs soft routing over tailors based on the semantics of condition instructions, enabling zero-shot generalization and parameter-efficient adaptation to novel conditions. To further improve condition fidelity and training efficiency, we introduce a representation alignment loss that aligns condition embeddings with early diffusion features. Extensive experiments demonstrate that DivControl achieves state-of-the-art controllability with 36.4$\times$ less training cost, while simultaneously improving average performance on basic conditions. It also delivers strong zero-shot and few-shot performance on unseen conditions, demonstrating superior scalability, modularity, and transferability.

图像生成控制生成模型压缩零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。