arXiv:2601.03382cs.CV2026-01被引 1

融合空间与频域特征,用血流检测提升深度伪造识别准确率。

A Novel Unified Approach to Deepfake Detection

  • 引入空间与频域交叉注意力,结合血流特征提取器
  • 在FF++和Celeb-DF上达99.88%和99.80% AUC
  • 模型跨数据集泛化能力强,适合实际部署场景

人工智能的发展带来了诸多威胁,其中最突出的是深度伪造的生成与滥用。为维护数字时代的信任,对深度伪造进行检测与标记至关重要。本文提出一种统一架构,用于图像与视频中的深度伪造检测。该架构利用空间与频域特征间的交叉注意力,并引入血流检测模块,以区分图像是否为真实。实验结果表明,采用Swin Transformer与BERT时,在FF++和Celeb-DF数据集上分别取得99.80%与99.88%的AUC;使用EfficientNet-B4与BERT时,对应AUC为99.55%与99.38%。该方法在跨数据集测试中表现优异,具备良好泛化能力。

原文摘要 · Abstract (English)

The advancements in the field of AI is increasingly giving rise to various threats. One of the most prominent of them is the synthesis and misuse of Deepfakes. To sustain trust in this digital age, detection and tagging of deepfakes is very necessary. In this paper, a novel architecture for Deepfake detection in images and videos is presented. The architecture uses cross attention between spatial and frequency domain features along with a blood detection module to classify an image as real or fake. This paper aims to develop a unified architecture and provide insights into each step. Though this approach we achieve results better than SOTA, specifically 99.80%, 99.88% AUC on FF++ and Celeb-DF upon using Swin Transformer and BERT and 99.55, 99.38 while using EfficientNet-B4 and BERT. The approach also generalizes very well achieving great cross dataset results as well.

深度伪造统一架构多模态检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。