arXiv:2606.11615cs.CVcs.CR2026-06被引 1

用文本控制生成假脸,骗过人脸识别系统

Adv-TGD: Adversarial Text-Guided Diffusion for Face Recognition Impersonation Attacks

论文配图:Adv-TGD: Adversarial Text-Guided Diffusion for Face Recognition Impersonation Attacks
图 1 · 摘自论文原文
  • 基于Stable Diffusion微调,用文本提示生成逼真伪造人脸
  • 攻击成功率85.9%,比现有方法最高提升16个百分点
  • 保留面部自然度,适合研究隐私漏洞与防御机制

面部识别技术的普及带来严重隐私风险,人脸数据可能在未经同意的情况下被滥用。为此,我们提出Adv-TGD,一种生成对抗性攻击框架,能够合成逼真的伪造人脸,欺骗面部识别系统。该框架基于Stable Diffusion v2.1,对每个源-目标身份对进行条件化文本提示的轻量级LoRA微调,在固定去噪步骤中优化跨注意力适配器。通过面部局部热力图掩码约束潜空间融合,实现精准的身份篡改并保留非敏感区域。设计复合目标函数,整合掩码ε-MSE重建、阈值化身份偏差、方向特征对齐及源相似性抑制,平衡攻击效果与视觉真实度。可选地,使用LLaVA生成的属性提示增强细粒度语义细节,且不引入身份线索。在黑盒评估下,平均攻击成功率(ASR)达85.90%,超越语义类SOTA方法Adv-CPG(+6.25)、扩散模型方法DiffAIM(+3)和噪声方法P3-Mask(+16)。尽管攻击性强,仍保持高视觉保真度(PSNR = 28.18 dB,SSIM = 0.981)。此外,框架可扩展至真实场景数据集(LADN)、通用图像分类(ImageNet)及Transformer扩散模型(FLUX.1)。

原文摘要 · Abstract (English)

The widespread adoption of face recognition (FR) technologies raises serious privacy concerns, as facial data can be exploited without consent. To address this challenge, we propose Adv-TGD, a generative adversarial attack framework that synthesizes photorealistic faces capable of impersonating target identities and deceiving face recognition systems. Built upon Stable Diffusion v2.1, Adv-TGD performs per-sample LoRA fine-tuning conditioned on concise textual prompts to generate natural yet adversarially manipulated identities. Unlike conventional identity attack approaches, our method optimizes lightweight cross-attention adapters for each source-target pair within a fixed-timestep denoising process. Latent blending is constrained by a face-local heatmap mask to ensure spatially precise identity manipulation while preserving non-sensitive regions. We introduce a composite objective that integrates masked epsilon-MSE reconstruction, thresholded identity divergence in FR embedding space, directional feature alignment, and source-similarity suppression to balance adversarial attack and visual realism. Optionally, LLaVA-generated attribute prompts enhance fine-grained semantic details without reintroducing identity cues. Under the black-box evaluation protocol, Adv-TGD attains an average attack success rate (ASR) of 85.90% across IR152, IRSE50, MobileFace, and FaceNet, surpassing the semantic SOTA baseline Adv-CPG by 6.25 points, the diffusion-based makeup method DiffAIM by 3 points, and the noise-based P3-Mask by 16 points. Despite its strong attack efficacy, Adv-TGD preserves high visual fidelity (PSNR = 28.18 dB, SSIM = 0.981). Furthermore, we demonstrate the flexibility of our framework by successfully extending it to in-the-wild datasets (LADN), general object classification (ImageNet), and transformer-based diffusion models (FLUX.1).

人脸伪造对抗攻击扩散模型隐私安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。