arXiv:2511.05219cs.CV2025-11NeurIPS被引 3

无需训练即可高效控制生成图像结构,一步提取注意力。

FreeControl: Efficient, Training-Free Structural Control via One-Step Attention Extraction

  • 仅从关键时间步提取一次注意力,复用至整个去噪过程。
  • 支持多源参考图组合控制,实现直观场景布局设计。
  • 相比原模型仅增加约5%开销,兼容主流扩散模型。

控制扩散模型生成图像的空间与语义结构仍是挑战。现有方法如ControlNet依赖手工构造条件图和重训练,限制了灵活性与泛化能力;基于反演的方法虽对齐更强,但因双路径去噪导致推理成本过高。本文提出FreeControl,一种无需训练的扩散模型语义结构控制框架。不同于以往在多个时间步提取注意力,FreeControl仅从一个最优选择的关键时间步执行单步注意力提取,并在整个去噪过程中复用。该方法实现高效结构引导,无需反演或重训练。为进一步提升质量和稳定性,提出潜在条件解耦(LCD):将关键时间步与去噪潜变量在注意力提取中分离,实现更精细的注意力控制并消除结构伪影。FreeControl还支持通过多源参考图进行组合控制,实现直观场景布局设计与更强提示对齐。该方法开创了测试时控制新范式,可直接从原始图像生成结构与语义一致、视觉连贯的结果,兼具直观组合设计能力与现代扩散模型兼容性,额外开销约为5%。

原文摘要 · Abstract (English)

Controlling the spatial and semantic structure of diffusion-generated images remains a challenge. Existing methods like ControlNet rely on handcrafted condition maps and retraining, limiting flexibility and generalization. Inversion-based approaches offer stronger alignment but incur high inference cost due to dual-path denoising. We present FreeControl, a training-free framework for semantic structural control in diffusion models. Unlike prior methods that extract attention across multiple timesteps, FreeControl performs one-step attention extraction from a single, optimally chosen key timestep and reuses it throughout denoising. This enables efficient structural guidance without inversion or retraining. To further improve quality and stability, we introduce Latent-Condition Decoupling (LCD): a principled separation of the key timestep and the noised latent used in attention extraction. LCD provides finer control over attention quality and eliminates structural artifacts. FreeControl also supports compositional control via reference images assembled from multiple sources - enabling intuitive scene layout design and stronger prompt alignment. FreeControl introduces a new paradigm for test-time control, enabling structurally and semantically aligned, visually coherent generation directly from raw images, with the flexibility for intuitive compositional design and compatibility with modern diffusion models at approximately 5 percent additional cost.

结构控制扩散模型零训练注意力提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。