提出统一检测人脸与全身伪造图像的框架,提升识别泛化能力。
On the Holistic Approach for Detecting Human Image Forgery
- 双分支结构:分别处理人脸区域与全身语义一致性
- 在多个数据集上达到领先检测性能,对多种伪造类型鲁棒
- 适合需要全面防范AI伪造内容的研究与安全应用
AI生成内容(AIGC)的快速发展加剧了深度伪造威胁,涵盖从面部操作到完整逼真人像合成。现有检测方法多局限于局部区域或特定类型,难以跨场景泛化。本文提出HuForDet,一种面向人像伪造的全貌检测框架,包含两个分支:(1) 人脸伪造检测分支,采用在RGB与频域运行的异构专家,含自适应拉普拉斯高斯(LoG)模块,可捕捉从细微拼接边界到粗粒度纹理异常的痕迹;(2) 上下文感知伪造检测分支,利用多模态大语言模型(MLLM)分析全身语义一致性,并通过置信度估计机制动态调整其在特征融合中的权重。我们构建了统合现有面部伪造数据与全新全身合成人像的新数据集HuFor。大量实验表明,HuForDet在多种人像伪造场景中均达到当前最优性能,并展现出更强的鲁棒性。
原文摘要 · Abstract (English)
The rapid advancement of AI-generated content (AIGC) has escalated the threat of deepfakes, from facial manipulations to the synthesis of entire photorealistic human bodies. However, existing detection methods remain fragmented, specializing either in facial-region forgeries or full-body synthetic images, and consequently fail to generalize across the full spectrum of human image manipulations. We introduce HuForDet, a holistic framework for human image forgery detection, which features a dual-branch architecture comprising: (1) a face forgery detection branch that employs heterogeneous experts operating in both RGB and frequency domains, including an adaptive Laplacian-of-Gaussian (LoG) module designed to capture artifacts ranging from fine-grained blending boundaries to coarse-scale texture irregularities; and (2) a contextualized forgery detection branch that leverages a Multi-Modal Large Language Model (MLLM) to analyze full-body semantic consistency, enhanced with a confidence estimation mechanism that dynamically weights its contribution during feature fusion. We curate a human image forgery (HuFor) dataset that unifies existing face forgery data with a new corpus of full-body synthetic humans. Extensive experiments show that our HuForDet achieves state-of-the-art forgery detection performance and superior robustness across diverse human image forgeries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。