arXiv:2508.01873cs.CV2025-08

用扩散模型同时检测伪造人脸并精确定位异常区域

DiffusionFF: A Diffusion-based Framework for Joint Face Forgery Detection and Fine-Grained Artifact Localization

  • 用预训练检测器作特征编码,扩散模型生成细粒度伪造定位图
  • 在多个数据集上达到当前最优检测性能,定位精度显著提升
  • 适合需要可解释性伪造检测的安防与内容审核场景

深度伪造技术的快速发展对鲁棒可靠的面部伪造检测算法提出更高要求。除了判断图像是否被篡改外,精确识别伪造痕迹的位置对提升模型可解释性和增强用户信任同样重要。为此,我们提出DiffusionFF,一种基于扩散模型的联合框架,实现面部伪造检测与细粒度伪造痕迹定位。核心思想是构建新型编码-解码结构:将预训练的伪造检测器作为强大的“伪造特征编码器”,并将去噪扩散模型重构为“伪造特征解码器”。在编码器提取的多尺度伪造相关特征条件下,解码器逐步合成精细的伪造定位图。随后,将该定位图与检测器中的高层语义特征融合,显著提升检测能力。大量实验表明,DiffusionFF在多个基准测试中均达到当前最优(SOTA)性能,验证了其卓越的有效性与可解释性。

原文摘要 · Abstract (English)

The rapid evolution of deepfake technologies demands robust and reliable face forgery detection algorithms. While determining whether an image has been manipulated remains essential, the ability to precisely localize forgery clues is also important for enhancing model explainability and building user trust. To address this dual challenge, we introduce DiffusionFF, a diffusion-based framework that simultaneously performs face forgery detection and fine-grained artifact localization. Our key idea is to establish a novel encoder-decoder architecture: a pretrained forgery detector serves as a powerful "artifact encoder", and a denoising diffusion model is repurposed as an "artifact decoder". Conditioned on multi-scale forgery-related features extracted by the encoder, the decoder progressively synthesizes a detailed artifact localization map. We then fuse this fine-grained localization map with high-level semantic features from the forgery detector, leading to substantial improvements in detection capability. Extensive experiments show that DiffusionFF achieves state-of-the-art (SOTA) performance across multiple benchmarks, underscoring its superior effectiveness and explainability.

伪造检测扩散模型定位分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。