arXiv:2511.07009cs.CV2025-11被引 1

旧数据训练的AI检测模型,半年后准确率下降超30%。

Performance Decay in Deepfake Detection: The Limitations of Training on Outdated Data

  • 用两阶段方法检测当前深伪内容,AUROC超99.8%
  • 模型在六个月内生成的新深伪内容上召回率下降超30%
  • 关键能力来自帧级静态特征,非时间不一致性

不断进步的深度伪造技术使恶意合成内容愈发难以与真实内容区分,加剧了虚假信息、欺诈和骚扰等威胁。我们提出一种简单有效的两阶段检测方法,在当代深伪内容上实现了超过99.8%的AUROC。然而,该高性能仅维持短暂时间:当评估使用六个月后生成技术制作的深伪内容时,模型召回率下降超过30%,表明性能随威胁演化显著衰减。分析揭示两个关键洞见:其一,持续性能依赖于持续收集大规模、多样化的数据集;其二,预测能力主要源于静态的帧级伪影,而非时间不一致性。因此,未来有效检测的关键在于快速数据采集与先进帧级特征探测器的开发。

原文摘要 · Abstract (English)

The continually advancing quality of deepfake technology exacerbates the threats of disinformation, fraud, and harassment by making maliciously-generated synthetic content increasingly difficult to distinguish from reality. We introduce a simple yet effective two-stage detection method that achieves an AUROC of over 99.8% on contemporary deepfakes. However, this high performance is short-lived. We show that models trained on this data suffer a recall drop of over 30% when evaluated on deepfakes created with generation techniques from just six months later, demonstrating significant decay as threats evolve. Our analysis reveals two key insights for robust detection. Firstly, continued performance requires the ongoing curation of large, diverse datasets. Second, predictive power comes primarily from static, frame-level artifacts, not temporal inconsistencies. The future of effective deepfake detection therefore depends on rapid data collection and the development of advanced frame-level feature detectors.

深伪检测性能衰减数据时效性帧级特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。