arXiv:2510.27475cs.CVcs.MM2025-10

用参考样本检测音视频伪造,提升跨数据集识别能力

Referee: Reference-aware Audiovisual Deepfake Detection

  • 以单张参考图构建身份锚点,捕捉说话人特有生物特征一致性
  • 在跨数据集和跨语言测试中达到99.4% AUC,优于现有方法
  • 适合需要泛化能力强的音频视频伪造检测场景

由先进生成模型产生的深度伪造内容正迅速带来严重威胁,但现有音视频深度伪造检测方法难以泛化到未见过的篡改手段。为此,我们提出一种新型参考感知音视频深度伪造检测方法——Referee,用于捕捉细微的身份差异。与过度依赖瞬时时空伪影的现有方法不同,Referee采用身份瓶颈和匹配模块,利用单次示例捕获的说话人特有线索作为生物特征锚点,建模其关系一致性。在FakeAVCeleb、FaceForensics++和KoDF上的大量实验表明,Referee在跨数据集和跨语言评估协议上均取得领先性能,包括在KoDF上达到99.4% AUC。结果表明,显式关联基于参考的生物特征先验是实现泛化且可靠的音视频取证的关键方向。代码已开源:https://github.com/ewha-mmai/referee。

原文摘要 · Abstract (English)

Deepfakes generated by advanced generative models have rapidly posed serious threats, yet existing audiovisual deepfake detection approaches struggle to generalize to unseen manipulation methods. To address this, we propose a novel reference-aware audiovisual deepfake detection method, called Referee to capture fine-grained identity discrepancies. Unlike existing methods that overfit to transient spatiotemporal artifacts, Referee employs identity bottleneck and matching modules to model the relational consistency of speaker-specific cues captured by a single one-shot example as a biometric anchor. Extensive experiments on FakeAVCeleb, FaceForensics++, and KoDF demonstrate that Referee achieves state-of-the-art results on cross-dataset and cross-language evaluation protocols, including a 99.4% AUC on KoDF. These results highlight that explicitly correlating reference-based biometric priors is a key frontier for achieving generalized and reliable audiovisual forensics. The code is available at https://github.com/ewha-mmai/referee.

深度伪造检测参考感知音视频取证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。