用3D CNN分析视频时间伪影,提升高质伪造视频检测能力
Deepfake Detection in Social Media: A Temporal Artifact Analysis Using 3D Convolutional Neural Networks

- 基于R3D-18的3D卷积网络,融合时间一致性正则化训练
- 128x128分辨率下跨数据集检测达76.4%,微调后更高
- 时间伪影比空间特征更稳定,适合社交媒体传播场景
合成面部视频在社交平台传播速度远超平台审核能力,加剧虚假信息与身份攻击风险。仅依赖帧级检测的方法在生成质量提升时性能显著下降:高质量128x128 GAN输出使纯空间检测准确率下降5个百分点,而时间不一致特征仍保留。本文提出基于R3D-18的3D卷积神经网络检测器,采用二元交叉熵与时间一致性正则化组合损失函数,在DeepfakeTIMIT数据集上处理16帧片段,并以Kinetics-400动作识别权重初始化。在128x128分辨率下,内部评估准确率达92.8%;未微调情况下迁移至FaceForensics++得76.4%,经少量微调后性能提升。消融实验表明,迁移学习贡献7.2个百分点,人脸追踪增加3.5个百分点,时间一致性正则化对高质量伪造品有额外增益。结果表明,时间伪影比空间特征更具泛化性,能抵抗社交平台重编码干扰。
原文摘要 · Abstract (English)
Synthetic facial videos have proliferated across social media faster than platform moderation can respond, raising the cost of disinformation and identity-based attacks. Frame-level deepfake detectors degrade sharply as generator quality increases; high-quality 128x128 GAN output cuts spatial-only accuracy by five percentage points while leaving temporal inconsistencies largely intact. We address this gap with a 3D Convolutional Neural Network detector based on R3D-18, trained with a composite loss that combines binary cross-entropy with a temporal-consistency regularizer. The model processes 16-frame clips from the DeepfakeTIMIT dataset and is initialized from Kinetics-400 action-recognition weights. We report 92.8% accuracy on intra-dataset evaluation at 128x128 resolution; cross-dataset transfer to FaceForensics++ without fine-tuning reaches 76.4%, rising after minimal fine-tuning. Ablation studies show that transfer learning contributes 7.2 percentage points and face tracking adds 3.5 points, while temporal consistency regularization provides additional gains on high-quality fakes. The results establish that temporal artifacts generalize more broadly than spatial ones, providing a detection signal that survives social-media re-encoding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。