通过增强特征不稳定性,实现隐蔽的自监督人脸攻击。
FIDA: Feature Instability-Driven Attack on Self-Supervised Facial Representation

- 设计特征不稳定性损失,让触发特征对扰动更敏感。
- 攻击成功率高,且在多种防御下仍有效。
- 适合研究安全性和鲁棒性的从业者参考。
自监督学习(SSL)模型易受后门攻击,但其在人脸表征中的系统性风险尚未受到足够关注。自监督人脸学习中身份特征的纠缠特性,给攻击隐蔽性带来独特挑战。为此,我们提出FIDA(特征不稳定性驱动攻击)框架。FIDA采用细微语义触发器进行注入,核心创新在于引入一种新型目标函数——特征不稳定性损失。该损失训练编码器,在攻击优化过程中沿采样的扰动方向增强触发特征的敏感性,从而避免传统攻击中固有的刚性特征模式。这使FIDA能有效规避现有基于扰动的防御机制。实验表明,FIDA在各类评估设置下均实现高攻击成功率,并普遍保持良性任务性能,对依赖人脸分析的真实多媒体应用构成显著威胁。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) models are vulnerable to backdoor attacks. However, the systemic risks they pose in face representation have received little attention. The entanglement of identity features in self-supervised face learning presents unique challenges for attack stealthiness. To address this gap, we propose FIDA (Feature Instability-Driven Attack), a novel backdoor attack framework. FIDA uses subtle semantic triggers for injection, but its key innovation is a novel objective called Feature Instability Loss. It trains the encoder to increase the sensitivity of triggered features along perturbation directions sampled during attack optimization . By preventing the backdoor from exhibiting the rigid feature patterns typical of previous attacks, FIDA effectively evades the evaluated perturbation-based defenses. Experiments show that FIDA achieves a high attack success rate and generally preserves benign utility across the evaluated settings , posing a significant threat to real-world multimedia applications relying on facial analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。