用对抗样本让语音无法控制人脸动画,保护隐私。
Silence is Golden: Leveraging Adversarial Examples to Nullify Audio Control in LDM-based Talking-Head Generation
- 设计双重损失,使音频无法操控人脸生成
- 在扩散模型中生成能抗净化的鲁棒干扰
- 适合关注AI伦理与深度伪造防御的研究者
基于潜在扩散模型(LDM)的人脸动画技术可生成高度逼真的同步视频,但易被滥用于诈骗、政治操纵和虚假信息。现有主动防御方法通过向人脸图像添加扰动来防护,但对音频驱动的视频生成无效,且扩散净化技术可轻易消除这些扰动。为此,我们提出Silencer,一种两阶段防护方法:首先引入消音损失,使音频信号无法控制生成过程;其次在LDM中加入反净化损失,优化反演隐空间特征以生成抗净化的鲁棒扰动。大量实验验证了Silencer在主动保护肖像隐私方面的有效性。本工作旨在唤起人工智能安全领域对人脸动画技术伦理问题的关注。代码已开源。
原文摘要 · Abstract (English)
Advances in talking-head animation based on Latent Diffusion Models (LDM) enable the creation of highly realistic, synchronized videos. These fabricated videos are indistinguishable from real ones, increasing the risk of potential misuse for scams, political manipulation, and misinformation. Hence, addressing these ethical concerns has become a pressing issue in AI security. Recent proactive defense studies focused on countering LDM-based models by adding perturbations to portraits. However, these methods are ineffective at protecting reference portraits from advanced image-to-video animation. The limitations are twofold: 1) they fail to prevent images from being manipulated by audio signals, and 2) diffusion-based purification techniques can effectively eliminate protective perturbations. To address these challenges, we propose Silencer, a two-stage method designed to proactively protect the privacy of portraits. First, a nullifying loss is proposed to ignore audio control in talking-head generation. Second, we apply anti-purification loss in LDM to optimize the inverted latent feature to generate robust perturbations. Extensive experiments demonstrate the effectiveness of Silencer in proactively protecting portrait privacy. We hope this work will raise awareness among the AI security community regarding critical ethical issues related to talking-head generation techniques. Code: https://github.com/yuangan/Silencer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。