arXiv:2607.13234cs.CVcs.AI2026-07

动态训练的深度伪造检测系统,显著提升真实场景下的检测能力。

Continuously Evolving Deepfake Detection: An Architecture and Public-Benchmark Evaluation of a Dynamic Detection System

  • 通过持续对抗竞赛更新训练数据,构建动态检测模型BMF
  • 在19个公开数据集上表现优异,视频检测AUC达0.936
  • 适合关注真实世界反伪造、需持续更新检测能力的研究者

当前学术基准上表现接近完美的深度伪造检测器,在真实场景中性能大幅下降:最新野外评估显示,顶尖开源模型的AUC下降45-50%。我们认为这一差距源于结构性问题:静态检测器仅一次训练,面对不断演进的生成技术。本文提出BitMind Forensics(BMF),基于Bittensor SN34开放对抗竞赛持续更新训练分布。我们在19个公开数据集上评估了历史版本模型,涵盖经典人脸替换套件(FaceForensics++、Celeb-DF v1/v2/++、DFDC、DFD、UADFV、DF40)及近期野外与AI生成媒体基准(Sumsub、Deepfake-Eval-2024、WildRF、Community Forensics、AIGCDetectBench、GenImage、AI-GenBench、AIGIBench、RAID、GenVidBench、GenVideo-100K)。BMF在Sumsub原始图像上达到0.936 AUC,全四条件操纵测试集(140万图像)平均AUC为0.872,抗扰动能力强(JPEG压缩下0.855,降采样下0.799),且经GPEN增强后检测率提升至0.996。在Deepfake-Eval-2024中,其图像检测性能(0.915)媲美最佳商用模型(0.90),视频检测(0.822)更优(商用0.79),远超最优开源模型(0.56和0.63)。在21个生成器的AI图像集上取得0.991 AUC,GenVidBench上达0.918;在DFDC(0.947 vs 0.843)和Celeb-DF v2(0.9985 vs 0.956)上均超越FF++训练模型,且在Celeb-DF++上统计无显著差异。时间序列分析显示,连续发布的旧版模型在未参与训练的生成器媒体上表现持续提升(图像从0.842升至0.902,视频从0.864升至0.936)。评估框架公开,论文发布时生产接口即为所评估快照,支持独立验证。

原文摘要 · Abstract (English)

Deepfake detectors that achieve near-perfect scores on academic benchmarks collapse on real-world content: recent in-the-wild evaluations report AUC drops of 45-50% for state-of-the-art open-source models. We argue this gap is structural: static detectors are trained once against a moving generative frontier. We present BitMind Forensics (BMF), trained through Bittensor SN34, an open adversarial competition that continually refreshes the training distribution. We evaluate one dated export comprising image, general-video, and human-video checkpoints across nineteen public datasets: the canonical face-swap suites (FaceForensics++, Celeb-DF v1/v2/++, DFDC, DFD, UADFV, DF40) and recent in-the-wild and AI-generated-media benchmarks (Sumsub, Deepfake-Eval-2024, WildRF, Community Forensics, AIGCDetectBench, GenImage, AI-GenBench, AIGIBench, RAID, GenVidBench, GenVideo-100K). BMF reaches 0.936 AUC on Sumsub's original images and 0.872 pooled AUC over its full four-condition manipulation battery (1.4M images), staying robust under perturbation (0.855 JPEG, 0.799 downscaled), while GPEN enhancement improves detection (0.996). On Deepfake-Eval-2024, it matches the best commercial detector on images (0.915 vs 0.90) and exceeds it on video (0.822 vs 0.79), far above the best open-source detectors (0.56 and 0.63). It reaches 0.991 AUC on a 21-generator AI-image panel and 0.918 on GenVidBench, and exceeds the FF++-trained frontier on DFDC (0.947 vs 0.843) and Celeb-DF v2 (0.9985 vs 0.956), both contamination-audited, with statistical parity on Celeb-DF++. In a temporal study, successive dated exports improve on held-out media from generators absent from the static baseline's training (image 0.842 to 0.902; video 0.864 to 0.936). Our evaluation harness is public, and at publication the production API serves the exact evaluated snapshot for independent verification.

深度伪造检测动态训练AUC提升真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。