arXiv:2603.28493cs.CV2026-03被引 2

通过分阶段控制生成过程,实现单张图概念解耦与精准编辑。

ConceptWeaver: Weaving Disentangled Concepts with Flow

  • 发现生成过程分三阶段:构图、实例化、细节优化。
  • 在实例化阶段注入语义偏移,实现高保真内容编辑。
  • 适合需要精细可控图像生成的设计师与研究人员。

预训练的流模型在合成复杂场景方面表现优异,但缺乏从单张真实图像中直接解耦和定制底层概念的机制。为此,我们提出一种新型微分探查技术,用于分析单个概念令牌对速度场的影响。研究发现,生成过程并非单一整体,而是分为三个阶段:初始的‘蓝图阶段’建立低频结构,随后的‘实例化阶段’内容概念达到峰值强度并自然解耦,形成最佳操作窗口,最后是不依赖概念的细节优化阶段。基于此发现,我们提出ConceptWeaver框架,通过阶段感知优化策略,从单张参考图像中学习特定概念的语义偏移,并利用创新的ConceptWeaver引导机制,在生成过程中适时注入这些偏移。大量实验表明,ConceptWeaver实现了高保真、可组合的合成与编辑,证明理解并利用流模型内在的阶段性特性是实现精确多粒度内容操控的关键。

原文摘要 · Abstract (English)

Pre-trained flow-based models excel at synthesizing complex scenes yet lack a direct mechanism for disentangling and customizing their underlying concepts from one-shot real-world sources. To demystify this process, we first introduce a novel differential probing technique to isolate and analyze the influence of individual concept tokens on the velocity field over time. This investigation yields a critical insight: the generative process is not monolithic but unfolds in three distinct stages. An initial \textbf{Blueprint Stage} establishes low-frequency structure, followed by a pivotal \textbf{Instantiation Stage} where content concepts emerge with peak intensity and become naturally disentangled, creating an optimal window for manipulation. A final concept-insensitive refinement stage then synthesizes fine-grained details. Guided by this discovery, we propose \textbf{ConceptWeaver}, a framework for one-shot concept disentanglement. ConceptWeaver learns concept-specific semantic offsets from a single reference image using a stage-aware optimization strategy that aligns with the three-stage framework. These learned offsets are then deployed during inference via our novel ConceptWeaver Guidance (CWG) mechanism, which strategically injects them at the appropriate generative stage. Extensive experiments validate that ConceptWeaver enables high-fidelity, compositional synthesis and editing, demonstrating that understanding and leveraging the intrinsic, staged nature of flow models is key to unlocking precise, multi-granularity content manipulation.

图像生成概念解耦流模型可控编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。