arXiv:2604.20317cs.CV2026-04

无需标注数据,用专家混合模型实现面部属性解耦编辑。

MD-Face: MoE-Enhanced Label-Free Disentangled Representation for Interactive Facial Attribute Editing

论文配图:MD-Face: MoE-Enhanced Label-Free Disentangled Representation for Interactive Facial Attribute Editing
图 1 · 摘自论文原文
  • 采用专家混合架构动态分配专家,提升语义向量独立性。
  • 引入几何感知损失,使语义向量与语义边界对齐,减少属性混淆。
  • 支持实时交互编辑,图像质量优于扩散模型,延迟更低。

基于GAN的面部属性编辑广泛应用于虚拟形象和社交媒体,但常因属性纠缠导致修改一个特征时意外影响其他特征。尽管监督式解耦表示学习可缓解此问题,却依赖大量标注数据,成本高昂。为此,我们提出MD-Face,一种基于混合专家(MoE)的无标签解耦表示学习框架。该框架采用带门控机制的MoE主干网络,动态分配专家,使模型能学习更具独立性的语义向量。为进一步降低属性纠缠,我们设计了一种几何感知损失,通过雅可比推前方法将每个语义向量与对应的语义边界向量(SBV)对齐。在ProGAN和StyleGAN上的实验表明,MD-Face优于无监督基线,并达到与有监督方法相当的性能。相比扩散模型,其生成图像质量更高、推理延迟更低,适用于交互式编辑场景。

原文摘要 · Abstract (English)

GAN-based facial attribute editing is widely used in virtual avatars and social media but often suffers from attribute entanglement, where modifying one face attribute unintentionally alters others. While supervised disentangled representation learning can address this, it relies heavily on labeled data, incurring high annotation costs. To address these challenges, we propose MD-Face, a label-free disentangled representation learning framework based on Mixture of Experts (MoE). MD-Face utilizes a MoE backbone with a gating mechanism that dynamically allocates experts, enabling the model to learn semantic vectors with greater independence. To further enhance attribute entanglement, we introduce a geometry-aware loss, which aligns each semantic vector with its corresponding Semantic Boundary Vector (SBV) through a Jacobian-based pushforward method. Experiments with ProGAN and StyleGAN show that MD-Face outperforms unsupervised baselines and competes with supervised ones. Compared to diffusion-based methods, it offers better image quality and lower inference latency, making it ideal for interactive editing.

面部编辑解耦表征MoE无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。