arXiv:2508.09847cs.CV2025-08

用对比嵌入和分割模型提升小数据下人脸生成的可控性

Enhancing Diffusion Face Generation with Contrastive Embeddings and SegFormer Guidance

  • 引入InfoNCE损失优化属性嵌入,增强语义对齐
  • 采用SegFormer编码器实现更精准的分割引导
  • 适合关注可控人脸生成的研究者与开发者

我们在小规模CelebAMask-HQ数据集上构建了扩散模型的人脸生成基准,评估了无条件与有条件生成流程。研究比较了UNet与DiT架构在无条件生成中的表现,并单独实验了基于LoRA微调预训练Stable Diffusion模型的方法。借鉴Giambi和Lisanti的多条件策略,结合属性向量与分割掩码,本文主要贡献在于引入InfoNCE损失优化属性嵌入,并采用SegFormer作为分割编码器。该改进显著提升了属性引导生成的语义一致性与可控性。结果表明,在数据有限的情况下,对比嵌入学习与先进分割编码对可控人脸生成具有显著效果。

原文摘要 · Abstract (English)

We present a benchmark of diffusion models for human face generation on a small-scale CelebAMask-HQ dataset, evaluating both unconditional and conditional pipelines. Our study compares UNet and DiT architectures for unconditional generation and explores LoRA-based fine-tuning of pretrained Stable Diffusion models as a separate experiment. Building on the multi-conditioning approach of Giambi and Lisanti, which uses both attribute vectors and segmentation masks, our main contribution is the integration of an InfoNCE loss for attribute embedding and the adoption of a SegFormer-based segmentation encoder. These enhancements improve the semantic alignment and controllability of attribute-guided synthesis. Our results highlight the effectiveness of contrastive embedding learning and advanced segmentation encoding for controlled face generation in limited data settings.

人脸生成扩散模型可控生成分割引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。