测试真实场景下深度伪造检测器效果,发现多数表现不佳。
Evaluating Deepfake Detectors in the Wild
- 构建超50万张高质量伪造图像数据集,模拟真实应用环境
- 超半数检测器AUC低于60%,最低仅50%
- 压缩、增强等简单操作可大幅降低检测性能,适合安全评估
由先进机器学习模型驱动的深度伪造对身份验证和数字媒体真实性构成重大且不断演变的威胁。尽管已开发出众多检测工具,但其在真实数据上的有效性尚未得到检验。本文评估现代深度伪造检测器,提出一种模拟真实场景的新型测试方法。利用最先进的生成技术,创建包含超过50万张高质量伪造图像的综合性数据集。分析表明,深度伪造检测仍具挑战性:在所测试的检测器中,不足一半AUC超过60%,最低为50%。研究还证明,JPEG压缩、图像增强等基础图像操作可显著降低模型性能。所有代码与数据已公开于https://github.com/SumSubstance/Deepfake-Detectors-in-the-Wild。
原文摘要 · Abstract (English)
Deepfakes powered by advanced machine learning models present a significant and evolving threat to identity verification and the authenticity of digital media. Although numerous detectors have been developed to address this problem, their effectiveness has yet to be tested when applied to real-world data. In this work we evaluate modern deepfake detectors, introducing a novel testing procedure designed to mimic real-world scenarios for deepfake detection. Using state-of-the-art deepfake generation methods, we create a comprehensive dataset containing more than 500,000 high-quality deepfake images. Our analysis shows that detecting deepfakes still remains a challenging task. The evaluation shows that in fewer than half of the deepfake detectors tested achieved an AUC score greater than 60%, with the lowest being 50%. We demonstrate that basic image manipulations, such as JPEG compression or image enhancement, can significantly reduce model performance. All code and data are publicly available at https://github.com/SumSubstance/Deepfake-Detectors-in-the-Wild.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。