数据多样性比频率不变性更能提升深伪检测的压缩鲁棒性
Data Diversity, Not Frequency Invariance: A Controlled and Self-Audited Study of Compression-Robust Deepfake Detection
- 通过控制训练配方和模型容量,验证压缩鲁棒性关键因素
- 真实多码率数据比合成JPEG增强使检测性能提升7.3 AUC点
- 频率特征融合在实验中未带来显著增益,反而可能降低效果
频率特征与压缩不变表示学习常被视为深伪检测在视频压缩下保持鲁棒性的关键。本文通过CAFRL(块DCT与FFT相位流、压缩级别条件化带注意力、对抗性梯度反转压缩不变性)进行受控测试。在预注册协议下,使用容量与增强匹配的对照组,纯EfficientNet-B0在多质量数据上于FaceForensics++测试集上,于每个压缩等级均优于CAFRL,CRF 40时领先3.66 AUC点(配对检验,单种子)。自审计发现四个偏向频率假设的缺陷,修复后基准模型反超相同配方版本3.96 AUC点。频率路径虽具判别力(独立验证AUC 0.91–0.98),但在融合框架中无边际贡献;在两种特征宽度、4.0M主干网络及所有种子池区间内均未显现差异。对抗分支按原设计无效且损害自身条件估计器;公平配方下未被测试。单次H.264重编码下的鲁棒性实源于数据多样性:真实恒定码率变体优于合成JPEG增强7.3点(单次运行,非重叠区间)。证据基于FaceForensics++家族、生成对抗网络时代、单一编码格式。需在训练配方与容量上严格匹配,应优先通过编码格式多样性获得鲁棒性,而非依赖架构改进。
原文摘要 · Abstract (English)
Frequency features and compression-invariant representation learning are widely assumed to be key to deepfake detection that survives video compression. We test this with CAFRL - block-DCT and FFT-phase streams, compression-level-conditioned band attention, and adversarial (gradient-reversal) compression invariance - and report a controlled negative. Under a pre-registered protocol with capacity- and augmentation-matched controls, a plain EfficientNet-B0 on multi-quality data beat CAFRL as specified at every compression level on the FaceForensics++ test split, by 3.66 AUC points at CRF 40 (paired, single seed). A self-audit of our own negative found four defects biased against the frequency hypothesis, and pre-specified re-tests repairing all four showed the deficit to be a recipe artifact, not an architecture failure: the baseline recipe recovered 3.96 points over the matching shipped-recipe variant. The frequency path made no detectable difference: discriminative alone (standalone validation AUC 0.91-0.98 late in training) but of no marginal value under this fusion, at two feature widths of one 4.0 M trunk, every seed-pooled interval for the intra-dataset compression contrasts including zero; on the single held-out manipulation tested, the fair variants sat below the plain backbone. The adversarial branch, as specified, added nothing and degraded its own conditioning estimator; at the fair recipe it is untested. Robustness under single-pass H.264 re-encoding came instead from data diversity: real constant-rate-factor variants beat synthetic JPEG augmentation by 7.3 points (single runs, non-overlapping intervals). The evidence is FaceForensics++-family, GAN-era and single-codec. Match controls on training recipe as well as capacity, and buy compression robustness with codec diversity before architecture.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。