arXiv:2607.21776cs.LGcs.CR2026-07

用生理信号检测说话脸深度伪造,突破传统图像检测盲区。

Physiological Signals as a Forensic Modality for Talking-Face Deepfake Detection

论文配图:Physiological Signals as a Forensic Modality for Talking-Face Deepfake Detection
图 1 · 摘自论文原文
  • 通过RhythmFormer提取视频级脉搏波形,仅依赖生理通道进行检测。
  • 在Celeb-DF++ TF子集上达AUC 0.806、EER 27.8%,接近最优通用模型。
  • 揭示不同生成方法的生理特征差异,为伪造类型提供可解释性依据。

说话脸(TF)深度伪造通过静态源图和音频合成逼真人脸视频,现有图像检测器普遍失效。由于该类伪造无真实视频基础,无法继承生理特征,远程光电容积脉搏波描记法(rPPG)成为独特检测模态。本文提出基于RhythmFormer提取每视频rPPG波形,并训练轻量级分类器区分真实与合成生理信号。在Celeb-DF++ TF子集上,采用严格主体无关协议(测试身份完全脱离训练集),1D ResNet达AUC 0.806、EER 27.8%,性能仅比最佳通用检测器(Effort, ICML 2025)低2.4点,且仅使用生理通道。我们复现了先前代表性rPPG检测器DeepFakesON-Phys,在旧式换脸数据上AUC 0.999,但在TF子集上降至0.622。进一步发现检测难度高度依赖生成方法:七种生成器间AUC范围从0.985(Real3DPortrait)到0.690(IP-LAP),且排名在所有评估协议中稳定,反映各生成器固有的可解释生理特性,构成主要理论贡献。

原文摘要 · Abstract (English)

Talking-face (TF) deepfake generation synthesizes photore- alistic facial video from a static source image and an au- dio signal, producing forgeries that current image-based detectors consistently fail to identify. Unlike face-swap ma- nipulation, TF synthesis has no underlying real video from which to inherit physiological characteristics, making re- mote photoplethysmography (rPPG) a uniquely motivated detection modality for this forgery category. We propose a detection framework that extracts per-video rPPG wave- forms via RhythmFormer and trains a suite of lightweight classifiers to distinguish real from synthesized physiologi- cal signals. Evaluated on the TF subset of Celeb-DF++ un- der a strict subject-independent protocol, where test identi- ties are completely separated from training identities, our 1D ResNet achieves an AUC of 0.806 and EER of 27.8%, placing it within 2.4 points of the best published general- purpose detector (Effort, ICML 2025) while operating ex- clusively on the physiological channel. We document a con- trolled reproduction study of DeepFakesON-Phys, the rep- resentative prior rPPG detector, demonstrating degrada- tion from AUC 0.999 on legacy face-swap data to 0.622 on the TF subset of Celeb-DF++. We further show that detec- tion difficulty is strongly method-dependent: AUC ranges from 0.985 (Real3DPortrait) to 0.690 (IP-LAP) across the seven TF generators, with the ranking remaining perfectly stable across all evaluation protocols. This spread reflects an interpretable physiological property of each generator rather than evaluation noise, and constitutes the primary theoretical contribution of the work.

深度伪造生理信号rPPG检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。