IdGlow无需掩码,通过动态调制身份实现多人图像生成的稳定与灵活平衡。
IdGlow: Dynamic Identity Modulation for Multi-Subject Generation
- 采用分阶段无掩码扩散模型,动态调节身份注入时机与强度。
- 在两人融合与年龄变换任务中,面部保真度与美学质量均达顶尖水平。
- 适合需要多身份协同生成的商业级图像创作场景。
多人图像生成需在统一场景中自然融合多个参考身份。现有方法依赖固定空间掩码或局部注意力,常陷入“稳定性-灵活性困境”,尤其在需复杂形变的任务(如保形年龄转换)中表现不佳。为此,我们提出IdGlow,一种基于流匹配扩散模型的无掩码、渐进式两阶段框架。在监督微调阶段,引入任务自适应时间步调度:线性衰减策略逐步放松约束以实现自然群体构图,时序门控机制将身份注入集中于关键语义窗口,成功保留成人面部特征而不覆盖儿童解剖结构。为解决属性泄漏与语义模糊问题,引入基于错误案例驱动的视觉-语言模型(VLM),实现精准、上下文感知的提示合成。第二阶段设计细粒度组级直接偏好优化(DPO),采用加权边界公式,同步消除多主体伪影、提升纹理和谐度,并校准身份保真度至真实分布。在两个挑战性基准——直接多人融合与年龄变换群体生成上,实验表明IdGlow从根本上缓解了稳定性-灵活性矛盾,在当前最优面部保真度与商业级美学质量之间实现了卓越权衡。
原文摘要 · Abstract (English)
Multi-subject image generation requires seamlessly harmonizing multiple reference identities within a coherent scene. However, existing methods relying on rigid spatial masks or localized attention often struggle with the "stability-plasticity dilemma," particularly failing in tasks that require complex structural deformations, such as identity-preserving age transformation. To address this, we present IdGlow, a mask-free, progressive two-stage framework built upon Flow Matching diffusion models. In the supervised fine-tuning (SFT) stage, we introduce task-adaptive timestep scheduling aligned with diffusion generative dynamics: a linear decay schedule that progressively relaxes constraints for natural group composition, and a temporal gating mechanism that concentrates identity injection within a critical semantic window, successfully preserving adult facial semantics without overriding child-like anatomical structures. To resolve attribute leakage and semantic ambiguity without explicit layout inputs, we further integrate a badcase-driven Vision-Language Model (VLM) for precise, context-aware prompt synthesis. In the second stage, we design a Fine-Grained Group-Level Direct Preference Optimization (DPO) with a weighted margin formulation to simultaneously eliminate multi-subject artifacts, elevate texture harmony, and recalibrate identity fidelity towards real-world distributions. Extensive experiments on two challenging benchmarks -- direct multi-person fusion and age-transformed group generation -- demonstrate that IdGlow fundamentally mitigates the stability-plasticity conflict, achieving a superior Pareto balance between state-of-the-art facial fidelity and commercial-grade aesthetic quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。