arXiv:2601.02359cs.CV2026-01被引 1

用音频生成表情的扩散模型,能零样本检测人脸伪造。

ExposeAnyone: Personalized Audio-to-Expression Diffusion Models Are Robust Zero-Shot Face Forgery Detectors

  • 基于自监督扩散模型,从音频生成表情序列并个性化特定人物。
  • 在多个数据集上比现有最优方法高4.22%平均AUC,且可检测Sora2生成视频。
  • 对模糊、压缩等真实场景干扰鲁棒,适合实际应用。

未知深度伪造检测仍是人脸伪造检测中最具挑战性的问题。现有最先进方法依赖已有伪造视频或伪伪造视频进行有监督训练,难以泛化到未见伪造类型,易过拟合特定伪造模式。相比之下,自监督方法更具泛化潜力,但现有工作难以仅通过自监督学习到判别性表征。本文提出ExposeAnyone,一种完全自监督的方法,基于扩散模型从音频生成表情序列。核心思想是:将模型个性化到特定个体后,通过扩散重建误差计算可疑视频与个性化个体的身份距离,实现目标人物的伪造检测。大量实验表明:1)在DF-TIMIT、DFDCP、KoDF和IDForge数据集上,平均AUC比之前最优方法提升4.22个百分点;2)能有效检测Sora2生成视频,而此前方法表现较差;3)对模糊、压缩等噪声具有高度鲁棒性,凸显其在真实场景中的适用性。

原文摘要 · Abstract (English)

Detecting unknown deepfake manipulations remains one of the most challenging problems in face forgery detection. Current state-of-the-art approaches fail to generalize to unseen manipulations, as they primarily rely on supervised training with existing deepfakes or pseudo-fakes, which leads to overfitting to specific forgery patterns. In contrast, self-supervised methods offer greater potential for generalization, but existing work struggles to learn discriminative representations only from self-supervision. In this paper, we propose ExposeAnyone, a fully self-supervised approach based on a diffusion model that generates expression sequences from audio. The key idea is, once the model is personalized to specific subjects using reference sets, it can compute the identity distances between suspected videos and personalized subjects via diffusion reconstruction errors, enabling person-of-interest face forgery detection. Extensive experiments demonstrate that 1) our method outperforms the previous state-of-the-art method by 4.22 percentage points in the average AUC on DF-TIMIT, DFDCP, KoDF, and IDForge datasets, 2) our model is also capable of detecting Sora2-generated videos, where the previous approaches perform poorly, and 3) our method is highly robust to corruptions such as blur and compression, highlighting the applicability in real-world face forgery detection.

人脸伪造检测扩散模型自监督零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。