arXiv:2601.13551cs.CV2026-01被引 1

构建百万级细粒度人脸伪造数据集,研究欺骗检测模型的真实漏洞

DiffFace-Edit: A Diffusion-Based Facial Dataset for Forgery-Semantic Driven Deepfake Detection Analysis

  • 基于扩散模型生成超200万张人脸伪造图像,覆盖8个面部区域的精细编辑
  • 发现真实与伪造样本混杂时,检测模型性能下降超过40%
  • 适合研究深度伪造检测、模型鲁棒性及对抗攻击的学者使用

生成模型如今能制造出难以察觉的精细伪造人脸,带来严重隐私风险。然而现有AI生成人脸数据集普遍缺乏对局部区域精细篡改样本的关注。此外,尚未有研究系统分析真实与伪造样本混杂(即“检测器逃避样本”)对检测器的实际影响。为此,我们提出DiffFace-Edit数据集:包含超过两百万张生成式伪造图像,涵盖眼睛、鼻子等八个面部区域的编辑,支持单区域与多区域组合编辑;并首次系统评估了检测器逃避样本对检测模型的影响。我们进一步提出跨域评估框架,结合IMDL方法进行综合分析。数据集将公开于https://github.com/ywh1093/DiffFace-Edit。

原文摘要 · Abstract (English)

Generative models now produce imperceptible, fine-grained manipulated faces, posing significant privacy risks. However, existing AI-generated face datasets generally lack focus on samples with fine-grained regional manipulations. Furthermore, no researchers have yet studied the real impact of splice attacks, which occur between real and manipulated samples, on detectors. We refer to these as detector-evasive samples. Based on this, we introduce the DiffFace-Edit dataset, which has the following advantages: 1) It contains over two million AI-generated fake images. 2) It features edits across eight facial regions (e.g., eyes, nose) and includes a richer variety of editing combinations, such as single-region and multi-region edits. Additionally, we specifically analyze the impact of detector-evasive samples on detection models. We conduct a comprehensive analysis of the dataset and propose a cross-domain evaluation that combines IMDL methods. Dataset will be available at https://github.com/ywh1093/DiffFace-Edit.

人脸伪造扩散模型检测鲁棒性数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。