arXiv:2509.00428cs.CV2025-09TPAMI被引 2

用专家混合机制提升人脸生成的可控性与真实感

Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation

  • 分区域设计全局与局部专家,分别处理整体结构和局部细节
  • 通过动态门控网络随扩散过程调整控制权重,实现精细调节
  • 支持零样本泛化,适合需要高可控性的生成与安全应用

可控人脸生成在生成建模中面临语义可控性与照片真实感之间平衡的挑战。现有方法难以将语义控制与生成流程解耦。本文基于扩散Transformer(DiT)的专家专业化潜力,提出Face-MoGLE框架:(1) 通过掩码条件空间分解实现语义解耦的潜在建模,支持精准属性操控;(2) 采用全局与局部专家混合结构,同时捕捉整体结构与区域级语义,实现细粒度控制;(3) 设计动态门控网络,生成随扩散步数和空间位置变化的时变系数。大量实验表明,该方法在多模态与单模态生成场景下均有效,具备强大的零样本泛化能力。项目主页见 https://github.com/XavierJiezou/Face-MoGLE。

原文摘要 · Abstract (English)

Controllable face generation poses critical challenges in generative modeling due to the intricate balance required between semantic controllability and photorealism. While existing approaches struggle with disentangling semantic controls from generation pipelines, we revisit the architectural potential of Diffusion Transformers (DiTs) through the lens of expert specialization. This paper introduces Face-MoGLE, a novel framework featuring: (1) Semantic-decoupled latent modeling through mask-conditioned space factorization, enabling precise attribute manipulation; (2) A mixture of global and local experts that captures holistic structure and region-level semantics for fine-grained controllability; (3) A dynamic gating network producing time-dependent coefficients that evolve with diffusion steps and spatial locations. Face-MoGLE provides a powerful and flexible solution for high-quality, controllable face generation, with strong potential in generative modeling and security applications. Extensive experiments demonstrate its effectiveness in multimodal and monomodal face generation settings and its robust zero-shot generalization capability. Project page is available at https://github.com/XavierJiezou/Face-MoGLE.

人脸生成扩散模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。