通过时空一致性与注意力机制,精准识别深度伪造视频的局部篡改痕迹。
Deepfake Detection with Spatio-Temporal Consistency and Attention
- 基于残差网络,融合空间注意力和时序距离注意力捕捉局部篡改特征。
- 在Deepfake-TIMIT和FaceForensics++数据集上准确率超越现有方法。
- 模型轻量高效,兼具检测精度与计算资源优势,适合实际部署。
深度伪造视频因日益逼真的表现引发广泛关注。自动化检测技术随之受到研究者重视。当前方法多依赖全局帧特征,忽视了伪造视频中潜在的时空不一致性,且难以关注到空间与时间维度上的细微、局部的模式变化。为此,我们提出一种新型神经网络检测器,在单帧及帧序列层面聚焦伪造视频的局部篡改特征。采用ResNet骨干网络,通过空间注意力机制强化浅层特征学习;空间分支进一步融合纹理增强的浅层特征与深层特征。同时,模型利用距离注意力机制处理帧序列,将时间注意力图与深层特征融合。整体模型以分类方式训练,用于检测伪造内容。在Deepfake-TIMIT和FaceForensics++两个主流大规模数据集上评估,性能显著优于现有方法。此外,本方法在内存占用和计算效率方面也优于同类技术。
原文摘要 · Abstract (English)
Deepfake videos are causing growing concerns among communities due to their ever-increasing realism. Naturally, automated detection of forged Deepfake videos is attracting a proportional amount of interest of researchers. Current methods for detecting forged videos mainly rely on global frame features and under-utilize the spatio-temporal inconsistencies found in the manipulated videos. Moreover, they fail to attend to manipulation-specific subtle and well-localized pattern variations along both spatial and temporal dimensions. Addressing these gaps, we propose a neural Deepfake detector that focuses on the localized manipulative signatures of the forged videos at individual frame level as well as frame sequence level. Using a ResNet backbone, it strengthens the shallow frame-level feature learning with a spatial attention mechanism. The spatial stream of the model is further helped by fusing texture enhanced shallow features with the deeper features. Simultaneously, the model processes frame sequences with a distance attention mechanism that further allows fusion of temporal attention maps with the learned features at the deeper layers. The overall model is trained to detect forged content as a classifier. We evaluate our method on two popular large data sets and achieve significant performance over the state-of-the-art methods.Moreover, our technique also provides memory and computational advantages over the competitive techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。