轻量级网络实时检测深度伪造,兼顾精度与速度
A Spatial-Frequency Aware Multi-Scale Fusion Network for Real-Time Deepfake Detection
- 融合空间纹理与频率特征的门控模块,捕捉细微篡改痕迹
- 多尺度特征交互高效,准确率高于现有方法且推理更快
- 适合视频会议、社交平台等实时场景部署
随着实时深度伪造生成技术的快速发展,伪造内容在视频会议和社交媒体中愈发逼真且泛滥。尽管现有先进检测器在标准基准上表现优异,但其计算开销大,难以在实际应用中实现实时部署。为此,我们提出一种轻量级但高效的实时深度伪造检测架构——空间-频率感知多尺度融合网络(SFMFNet)。设计了空间-频率混合感知模块,通过门控机制联合利用空间纹理与频率伪影,提升对细微篡改的敏感性;引入令牌选择性交叉注意力机制,实现高效多层级特征交互;并采用残差增强的模糊池化结构,在下采样过程中保留关键语义信息。在多个基准数据集上的实验表明,SFMFNet在精度与效率之间取得了良好平衡,具备强泛化能力,具有实际应用价值。
原文摘要 · Abstract (English)
With the rapid advancement of real-time deepfake generation techniques, forged content is becoming increasingly realistic and widespread across applications like video conferencing and social media. Although state-of-the-art detectors achieve high accuracy on standard benchmarks, their heavy computational cost hinders real-time deployment in practical applications. To address this, we propose the Spatial-Frequency Aware Multi-Scale Fusion Network (SFMFNet), a lightweight yet effective architecture for real-time deepfake detection. We design a spatial-frequency hybrid aware module that jointly leverages spatial textures and frequency artifacts through a gated mechanism, enhancing sensitivity to subtle manipulations. A token-selective cross attention mechanism enables efficient multi-level feature interaction, while a residual-enhanced blur pooling structure helps retain key semantic cues during downsampling. Experiments on several benchmark datasets show that SFMFNet achieves a favorable balance between accuracy and efficiency, with strong generalization and practical value for real-time applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。