融合四类模型特征,提升多维度人脸伪造检测精度
Hierarchical Deep Fusion Framework for Multi-dimensional Facial Forgery Detection -- The 2024 Global Deepfake Image Detection Challenge
- 分层融合四个预训练模型的特征表示
- 在竞赛私有榜单上取得0.96852的得分
- 适合需要高精度伪造检测的安防场景
深度伪造技术的泛滥对数字安全与真实性构成严重挑战。针对多种操纵手法的检测,需依赖鲁棒且泛化的模型。本文提出分层深度融合框架(HDFF),一种基于集成学习的深度学习架构,用于高性能人脸伪造检测。该框架整合了Swin-MLP、CoAtNet、EfficientNetV2和DaViT四种不同预训练子模型,通过在MultiFFDI数据集上的多阶段微调,提取其专长特征。将各模型输出的特征向量拼接后,输入最终分类器进行训练,有效融合各模型优势。该方法在竞赛私有榜单上获得0.96852的得分,位列184支参赛队伍中的第20名,验证了分层融合在复杂图像分类任务中的有效性。
原文摘要 · Abstract (English)
The proliferation of sophisticated deepfake technology poses significant challenges to digital security and authenticity. Detecting these forgeries, especially across a wide spectrum of manipulation techniques, requires robust and generalized models. This paper introduces the Hierarchical Deep Fusion Framework (HDFF), an ensemble-based deep learning architecture designed for high-performance facial forgery detection. Our framework integrates four diverse pre-trained sub-models, Swin-MLP, CoAtNet, EfficientNetV2, and DaViT, which are meticulously fine-tuned through a multi-stage process on the MultiFFDI dataset. By concatenating the feature representations from these specialized models and training a final classifier layer, HDFF effectively leverages their collective strengths. This approach achieved a final score of 0.96852 on the competition's private leaderboard, securing the 20th position out of 184 teams, demonstrating the efficacy of hierarchical fusion for complex image classification tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。