arXiv:2512.15433cs.CV2025-12AAAI被引 4

用CLIP提升人脸模板逆向生成的细节真实度和攻击迁移性。

CLIP-FTI: Fine-Grained Face Template Inversion via CLIP-Driven Attribute Conditioning

  • 通过CLIP提取面部属性语义,融合到生成模型中增强细节。
  • 重建图像在身份识别和属性相似度上均优于现有方法。
  • 首次引入额外语义信息实现高精度逆向,适合隐私安全研究者。

人脸识别系统存储的人脸模板一旦泄露,可被逆向生成逼真替身,威胁隐私与身份安全。现有方法虽能生成较真实图像,但面部局部特征(如眼、鼻、口)过度平滑,且跨模型攻击迁移能力弱。本文提出CLIP-FTI,一种基于CLIP驱动的细粒度属性条件化人脸模板逆向框架。核心思路是利用CLIP模型获取面部特征的语义嵌入,通过跨模态特征交互网络将这些嵌入与泄露模板融合,并投影至预训练StyleGAN的中间潜在空间,由生成器合成与原始模板同身份但具有更精细面部属性的图像。在多个主流人脸识别骨干网络和数据集上的实验表明,该方法在(i)身份识别准确率与属性相似度上表现更优,(ii)恢复了更清晰的部件级属性语义,(iii)显著提升跨模型攻击迁移能力。据我们所知,这是首个在人脸模板逆向中引入模板外信息并达到当前最优性能的方法。

原文摘要 · Abstract (English)

Face recognition systems store face templates for efficient matching. Once leaked, these templates pose a threat: inverting them can yield photorealistic surrogates that compromise privacy and enable impersonation. Although existing research has achieved relatively realistic face template inversion, the reconstructed facial images exhibit over-smoothed facial-part attributes (eyes, nose, mouth) and limited transferability. To address this problem, we present CLIP-FTI, a CLIP-driven fine-grained attribute conditioning framework for face template inversion. Our core idea is to use the CLIP model to obtain the semantic embeddings of facial features, in order to realize the reconstruction of specific facial feature attributes. Specifically, facial feature attribute embeddings extracted from CLIP are fused with the leaked template via a cross-modal feature interaction network and projected into the intermediate latent space of a pretrained StyleGAN. The StyleGAN generator then synthesizes face images with the same identity as the templates but with more fine-grained facial feature attributes. Experiments across multiple face recognition backbones and datasets show that our reconstructions (i) achieve higher identification accuracy and attribute similarity, (ii) recover sharper component-level attribute semantics, and (iii) improve cross-model attack transferability compared to prior reconstruction attacks. To the best of our knowledge, ours is the first method to use additional information besides the face template attack to realize face template inversion and obtains SOTA results.

人脸逆向生成对抗网络隐私安全CLIP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。