arXiv:2510.16320cs.CV2025-10被引 7

发现深度伪造检测性能随数据规模呈幂律增长,可预测未来所需数据量。

Scaling Laws for Deepfake Detection

  • 通过构建超大规模数据集 ScaleDF,研究检测模型在真实图像域与伪造方法数量上的扩展规律。
  • 检测误差随真实域或伪造方法增加呈幂律下降,且可预测达到目标性能所需的数据量。
  • 强调数据驱动应对伪造技术演进,适合关注深度伪造防御与数据策略的研究者。

本文系统研究了深度伪造检测任务的缩放规律。我们分析了模型性能随真实图像域数量、深度伪造生成方法数量及训练图像数量的变化关系。由于现有数据集无法满足研究需求,我们构建了目前该领域最大的数据集 ScaleDF,包含超过 580 万张来自 51 个不同数据集(域)的真实图像,以及超过 880 万张由 102 种深度伪造方法生成的虚假图像。利用 ScaleDF,我们观察到检测误差随真实域或伪造方法数量增加呈现类似大语言模型的幂律衰减。这一关键发现不仅可用于预测达到目标性能所需的额外真实域或伪造方法数量,还启发我们以数据为中心应对不断演进的伪造技术。此外,我们考察了预训练和数据增强在缩放条件下的作用,以及缩放本身的局限性。

原文摘要 · Abstract (English)

This paper presents a systematic study of scaling laws for the deepfake detection task. Specifically, we analyze the model performance against the number of real image domains, deepfake generation methods, and training images. Since no existing dataset meets the scale requirements for this research, we construct ScaleDF, the largest dataset to date in this field, which contains over 5.8 million real images from 51 different datasets (domains) and more than 8.8 million fake images generated by 102 deepfake methods. Using ScaleDF, we observe power-law scaling similar to that shown in large language models (LLMs). Specifically, the average detection error follows a predictable power-law decay as either the number of real domains or the number of deepfake methods increases. This key observation not only allows us to forecast the number of additional real domains or deepfake methods required to reach a target performance, but also inspires us to counter the evolving deepfake technology in a data-centric manner. Beyond this, we examine the role of pre-training and data augmentations in deepfake detection under scaling, as well as the limitations of scaling itself.

深度伪造数据规模检测模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。