arXiv:2602.00627cs.CV2026-02中稿 · ICANN 2025被引 1

仅用一张图生成高保真人脸,无需调参即可定制肖像。

FaceSnap: Enhanced ID-fidelity Network for Tuning-free Portrait Customization

  • 通过融合低层细节与高层语义特征,实现精准人脸信息提取。
  • 单次推理生成一致人脸,身份保持率显著优于现有方法。
  • 适配多类扩散模型,适合快速定制肖像的实用场景。

得益于文本到图像扩散模型的重大进展,个性化图像生成尤其是定制化肖像生成近年来取得显著进步。然而,现有方法或需耗时微调且泛化能力差,或难以保持面部细节高保真度。为此,我们提出FaceSnap,一种基于Stable Diffusion(SD)的新方法,仅需单张参考图,即可在单次推理中生成高度一致的结果。该方法即插即用,可轻松扩展至不同SD模型。具体而言,设计了新型面部属性混合器(Facial Attribute Mixer),从低层特定特征与高层抽象特征中提取综合融合信息,为图像生成提供更好引导;引入关键点预测器(Landmark Predictor),在不同姿态下维持参考身份一致性,为生成提供多样且精细的空间控制条件;最后通过身份保持模块将这些信息注入UNet。实验表明,本方法在个性化定制肖像生成方面表现卓越,显著超越当前最优方法。

原文摘要 · Abstract (English)

Benefiting from the significant advancements in text-to-image diffusion models, research in personalized image generation, particularly customized portrait generation, has also made great strides recently. However, existing methods either require time-consuming fine-tuning and lack generalizability or fail to achieve high fidelity in facial details. To address these issues, we propose FaceSnap, a novel method based on Stable Diffusion (SD) that requires only a single reference image and produces extremely consistent results in a single inference stage. This method is plug-and-play and can be easily extended to different SD models. Specifically, we design a new Facial Attribute Mixer that can extract comprehensive fused information from both low-level specific features and high-level abstract features, providing better guidance for image generation. We also introduce a Landmark Predictor that maintains reference identity across landmarks with different poses, providing diverse yet detailed spatial control conditions for image generation. Then we use an ID-preserving module to inject these into the UNet. Experimental results demonstrate that our approach performs remarkably in personalized and customized portrait generation, surpassing other state-of-the-art methods in this domain.

人脸生成扩散模型个性化定制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。