arXiv:2502.07269cs.CV2025-02被引 2

自动选数据提升深度伪造检测模型持续学习能力

Exploring Active Data Selection Strategies for Continuous Training in Deepfake Detection

  • 用检测模型置信度从冗余数据池中主动筛选新训练样本
  • 仅用15%数据量即实现2.5%的错误率,性能显著提升
  • 适合需频繁更新的深度伪造检测系统使用

在深度伪造检测中,随着新型伪造方法不断出现,必须持续调整检测模型参数以维持高性能。本文提出一种自动且主动的数据选择方法,用于在模型定期更新时,从包含大量新伪造图像和真实图像的冗余数据池中,选取少量新增训练数据进行连续训练。该方法以检测模型的置信度作为筛选指标,自动挑选最具价值的数据。实验表明,仅使用数据池中15%的样本进行持续训练,检测模型的性能显著提升,错误率(EER)降至2.5%,验证了该策略在小样本条件下的高效性与有效性。

原文摘要 · Abstract (English)

In deepfake detection, it is essential to maintain high performance by adjusting the parameters of the detector as new deepfake methods emerge. In this paper, we propose a method to automatically and actively select the small amount of additional data required for the continuous training of deepfake detection models in situations where deepfake detection models are regularly updated. The proposed method automatically selects new training data from a \textit{redundant} pool set containing a large number of images generated by new deepfake methods and real images, using the confidence score of the deepfake detection model as a metric. Experimental results show that the deepfake detection model, continuously trained with a small amount of additional data automatically selected and added to the original training set, significantly and efficiently improved the detection performance, achieving an EER of 2.5% with only 15% of the amount of data in the pool set.

深度伪造主动学习持续学习数据筛选

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。