arXiv:2602.00639cs.CV2026-02中稿 · Information Fusion…被引 7

零样本生成人脸,保身份还控表情姿势

Diff-PC: Identity-preserving and 3D-aware Controllable Diffusion for Zero-shot Portrait Customization

  • 用3D人脸重建+双编码器融合,精准保留身份特征
  • 在多个数据集上身份相似度达90.1%,表情控制更自然
  • 适合需要高保真人脸定制的影视/游戏应用

肖像定制(PC)因广泛应用前景受到关注。现有方法在身份(ID)保持和面部控制方面存在不足。为此,我们提出Diff-PC,一种基于扩散模型的零样本肖像定制框架,可生成高身份保真度、指定面部属性及多样化背景的逼真肖像。具体而言,该方法利用3D人脸预测器重建包含参考身份、目标表情与姿态的3D-aware面部先验。为捕捉精细面部细节,设计了融合局部与全局特征的ID-Encoder。随后,提出基于3D人脸的ID-Ctrl以对齐身份特征。进一步引入ID-Injector提升身份保真度与面部可控性。在自建的以身份为中心的数据集上训练后,显著提升了面部相似度与文本到图像(T2I)一致性。大量实验表明,Diff-PC在身份保持、面部控制与T2I一致性方面均优于现有方法。此外,本方法兼容多风格基础模型。

原文摘要 · Abstract (English)

Portrait customization (PC) has recently garnered significant attention due to its potential applications. However, existing PC methods lack precise identity (ID) preservation and face control. To address these tissues, we propose Diff-PC, a diffusion-based framework for zero-shot PC, which generates realistic portraits with high ID fidelity, specified facial attributes, and diverse backgrounds. Specifically, our approach employs the 3D face predictor to reconstruct the 3D-aware facial priors encompassing the reference ID, target expressions, and poses. To capture fine-grained face details, we design ID-Encoder that fuses local and global facial features. Subsequently, we devise ID-Ctrl using the 3D face to guide the alignment of ID features. We further introduce ID-Injector to enhance ID fidelity and facial controllability. Finally, training on our collected ID-centric dataset improves face similarity and text-to-image (T2I) alignment. Extensive experiments demonstrate that Diff-PC surpasses state-of-the-art methods in ID preservation, facial control, and T2I consistency. Furthermore, our method is compatible with multi-style foundation models.

人脸生成扩散模型身份保持零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。