通过分层特征捕捉面部区域依赖关系,提升深伪检测泛化能力
Face2Parts: Exploring Coarse-to-Fine Inter-Regional Facial Dependencies for Generalized Deepfake Detection
- 分阶段提取全图、人脸及关键部位特征,构建粗到细的层级表示
- 在多个数据集上平均AUC达98.42%,部分数据集表现接近100%
- 适合需要高泛化能力的深伪检测场景,尤其对抗新型伪造手法
多媒体数据在监控、人机交互、生物识别、证据采集和广告等领域广泛应用,但伪造者可制作深伪内容用于恶意目的。现有取证方法虽能检测特定面部区域的伪造痕迹,如面部轮廓、嘴唇、眼睛或鼻子,但泛化能力受限。本文提出一种名为Face2Parts的混合方法,基于分层特征表示(HFR),通过分阶段提取图像帧、人脸及关键区域(唇、眼、鼻)特征,利用通道注意力机制与深度三元组学习捕捉面部区域间的相互依赖关系。在多个基准数据集上评估,该方法在FF++上平均AUC达98.42%,CDF1为79.80%,CDF2为85.34%,DFD为89.41%,DFDC为84.07%,DTIM为95.62%,PDD为80.76%,WLDR为100%。结果表明,该方法具有强泛化性能,优于现有方法。
原文摘要 · Abstract (English)
Multimedia data, particularly images and videos, is integral to various applications, including surveillance, visual interaction, biometrics, evidence gathering, and advertising. However, amateur or skilled counterfeiters can simulate them to create deepfakes, often for slanderous motives. To address this challenge, several forensic methods have been developed to ensure the authenticity of the content. The effectiveness of these methods depends on their focus, with challenges arising from the diverse nature of manipulations. In this article, we analyze existing forensic methods and observe that each method has unique strengths in detecting deepfake traces by focusing on specific facial regions, such as the frame, face, lips, eyes, or nose. Considering these insights, we propose a novel hybrid approach called Face2Parts based on hierarchical feature representation ($HFR$) that takes advantage of coarse-to-fine information to improve deepfake detection. The proposed method involves extracting features from the frame, face, and key facial regions (i.e., lips, eyes, and nose) separately to explore the coarse-to-fine relationships. This approach enables us to capture inter-dependencies among facial regions using a channel-attention mechanism and deep triplet learning. We evaluated the proposed method on benchmark deepfake datasets in both intra-, inter-dataset, and inter-manipulation settings. The proposed method achieves an average AUC of 98.42\% on FF++, 79.80\% on CDF1, 85.34\% on CDF2, 89.41\% on DFD, 84.07\% on DFDC, 95.62\% on DTIM, 80.76\% on PDD, and 100\% on WLDR, respectively. The results demonstrate that our approach generalizes effectively and achieves promising performance to outperform the existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。