分离主体与场景引导路径,解决图像生成中身份保真与场景编辑的冲突。
Decoupled Guidance: Disentangling Subject and Context Pathways in Text-to-Image Personalization

- 通过双独立引导流分离主体身份与场景上下文
- 实现身份保真度提升15%且场景适配性更强
- 可插拔适配多种模型,支持推理时调节保真与编辑平衡
文本到图像个性化旨在将用户提供的主体生成于由文本描述的新场景中。然而,现有方法通常通过同一条件路径同时编码主体身份(保真度)和场景上下文(可编辑性),导致两者争夺注意力图资源,形成保真度-可编辑性权衡。本文称之为条件纠缠,并通过替换目标主体标记为通用标记,验证其对注意力分配和上下文遵循的影响。为此,提出一种即插即用的解耦引导(DeGu)框架,将主体身份与场景上下文分别通过两条独立引导流处理。进一步引入空间混合机制,动态融合两路信号,确保每路在语义相关区域运行而不互相干扰。DeGu 可直接应用于现有个性化方法,无需修改底层骨干模型,在多种方法与骨干网络(包括 flow-matching Diffusion Transformers, DiTs)上持续提升整体个性化性能,并支持推理时对保真度与可编辑性比例的灵活控制。
原文摘要 · Abstract (English)
Text-to-image personalization aims to generate a user-provided subject in novel scenes described by text. However, most existing methods encode subject identity (fidelity) and context (editability) through the same conditioning pathway, forcing the two to compete for attention-map resources. We refer to this phenomenon as conditioning entanglement and show that it induces a fidelity-editability trade-off. We further provide causal evidence by replacing the target subject token with a generic subject token, which produces shifts in attention allocation and corresponding changes in context adherence. To this end, we propose Decoupled Guidance (DeGu), a plug-and-play framework that routes subject identity and scene context through two independent guidance streams. We further introduce a spatial mixing mechanism that dynamically fuses these streams, ensuring each operates within its semantically relevant region without interference. Furthermore, DeGu can be readily applied to existing personalization methods without modifying the underlying backbone models, consistently improving the overall personalization performance while enabling inference-time control over the fidelity-editability balance, across diverse methods and backbones, including flow-matching Diffusion Transformers (DiTs).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。