arXiv:2505.11110cs.CV2025-05被引 5

通过分析生成图像的频域特征,可精准追溯其训练数据集来源。

Deepfake Forensic Analysis: Source Dataset Attribution and Legal Implications of Synthetic Media Manipulation

  • 结合频域变换与颜色分布,提取合成图像中的数据集特异性痕迹。
  • 在多种GAN模型上实现98%-99%的真假图像分类与数据集归属准确率。
  • 为版权保护、隐私合规和内容治理提供可落地的数字取证工具。

由生成对抗网络(GAN)生成的合成媒体在真实性验证与数据源追溯方面带来严峻挑战,引发版权保护、隐私安全及法律合规的广泛关注。本文提出一种新型数字取证框架,通过可解释特征分析,识别GAN生成图像所源自的训练数据集(如CelebA或FFHQ)。该方法融合傅里叶/离散余弦变换(DCT/FFT)、颜色分布度量及局部特征描述符(SIFT),提取合成输出中嵌入的判别性统计特征。监督分类器(随机森林、SVM、XGBoost)在多种GAN架构(StyleGAN、AttGAN、GDWCT、StarGAN、StyleGAN2)下,实现98%-99%的二分类(真实/合成)与多类数据集归属准确率。实验表明,频域特征(特别是DCT/FFT)能有效捕捉特定数据集的重采样模式与频谱异常,而颜色直方图则揭示了生成过程中的隐式正则化策略。此外,本文探讨了数据集溯源在应对版权侵权、个人数据滥用及符合GDPR与加州AB 602法案方面的法律意义。该框架推动生成模型的问责与治理,适用于数字取证、内容审核与知识产权诉讼。

原文摘要 · Abstract (English)

Synthetic media generated by Generative Adversarial Networks (GANs) pose significant challenges in verifying authenticity and tracing dataset origins, raising critical concerns in copyright enforcement, privacy protection, and legal compliance. This paper introduces a novel forensic framework for identifying the training dataset (e.g., CelebA or FFHQ) of GAN-generated images through interpretable feature analysis. By integrating spectral transforms (Fourier/DCT), color distribution metrics, and local feature descriptors (SIFT), our pipeline extracts discriminative statistical signatures embedded in synthetic outputs. Supervised classifiers (Random Forest, SVM, XGBoost) achieve 98-99% accuracy in binary classification (real vs. synthetic) and multi-class dataset attribution across diverse GAN architectures (StyleGAN, AttGAN, GDWCT, StarGAN, and StyleGAN2). Experimental results highlight the dominance of frequency-domain features (DCT/FFT) in capturing dataset-specific artifacts, such as upsampling patterns and spectral irregularities, while color histograms reveal implicit regularization strategies in GAN training. We further examine legal and ethical implications, showing how dataset attribution can address copyright infringement, unauthorized use of personal data, and regulatory compliance under frameworks like GDPR and California's AB 602. Our framework advances accountability and governance in generative modeling, with applications in digital forensics, content moderation, and intellectual property litigation.

深度伪造数字取证GAN溯源版权合规

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。