识别换脸伪造视频的生成模型,快速准确且轻量级。
FAME: A Lightweight Spatio-Temporal Network for Model Attribution of Face-Swap Deepfakes
- 融合时空注意力机制,捕捉不同换脸模型的细微痕迹。
- 在三个数据集上准确率领先,推理速度更快。
- 适合实际部署的数字取证与信息安全场景。
换脸深度伪造视频的广泛出现对数字安全、隐私和媒体真实性构成日益严重的威胁,亟需有效的溯源工具以识别伪造来源。尽管以往研究多集中于二分类的深度伪造检测,但模型归属(即判断某段伪造视频由哪个生成模型制作)仍鲜有探索。本文提出FAME(Fake Attribution via Multilevel Embeddings),一种轻量级高效的时空框架,用于捕捉不同换脸模型特有的细微生成痕迹。FAME结合空间与时间注意力机制,在保持计算高效的同时提升归属准确性。我们在三个具有挑战性且多样化的数据集——Deepfake Detection and Manipulation(DFDM)、FaceForensics++ 和 FakeAVCeleb——上进行评估。结果表明,FAME在准确率与运行时间上均显著优于现有方法,展现出在真实世界数字取证与信息安全应用中的巨大潜力。
原文摘要 · Abstract (English)
The widespread emergence of face-swap Deepfake videos poses growing risks to digital security, privacy, and media integrity, necessitating effective forensic tools for identifying the source of such manipulations. Although most prior research has focused primarily on binary Deepfake detection, the task of model attribution -- determining which generative model produced a given Deepfake -- remains underexplored. In this paper, we introduce FAME (Fake Attribution via Multilevel Embeddings), a lightweight and efficient spatio-temporal framework designed to capture subtle generative artifacts specific to different face-swap models. FAME integrates spatial and temporal attention mechanisms to improve attribution accuracy while remaining computationally efficient. We evaluate our model on three challenging and diverse datasets: Deepfake Detection and Manipulation (DFDM), FaceForensics++, and FakeAVCeleb. Results show that FAME consistently outperforms existing methods in both accuracy and runtime, highlighting its potential for deployment in real-world forensic and information security applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。