通过物理规律引导注意力,提升深度伪造视频检测的泛化与抗干扰能力
Aletheia: Physics-Conditioned Localized Artifact Attention (PhyLAA-X) for End-to-End Generalizable and Robust Deepfake Video Detection

- 将光学流动、反光不一致和心率信号等物理特征融入注意力机制
- 在跨生成器场景下准确率达97.2%(FaceForensics++),对抗攻击下仍保持79.4%准确率
- 适合需要高鲁棒性检测的安防、媒体审核场景
当前最先进的深度伪造检测器在同源数据上表现接近完美,但在跨生成器迁移、强压缩和对抗扰动下性能显著下降。核心问题在于语义伪影学习与物理规律解耦:如光流不连续、镜面反射不一致及心率调制反射(rPPG)未被有效建模。本文提出PhyLAA-X,是局部伪影注意力(LAA-X)的物理条件化扩展。该方法通过交叉注意力门控与共振一致性损失,将三个可微分的物理特征体——光流旋度、镜面反射偏度、空间上采样的rPPG功率谱——直接注入LAA-X注意力计算中。这迫使网络聚焦于语义不一致与物理违规共现的篡改边界,而这类区域生成模型难以持续复现。PhyLAA-X嵌入高效时空集成框架(EfficientNet-B4+BiLSTM, ResNeXt-101+Transformer, Xception+causal Conv1D),并采用不确定性感知自适应加权。在FaceForensics++(c23)上达到97.2%准确率/0.992 AUC-ROC;Celeb-DF v2上为94.9%/0.981;DFDC上为90.8%/0.966,优于最强基线(LAA-Net [1])4.1–7.3%。在ε=0.02的PGD-10攻击下仍保持79.4%准确率。单骨干消融实验表明,PhyLAA-X本身即可带来4.2%跨数据集AUC提升。完整系统已开源(https://github.com/devghori1264/Aletheia v1.2,2026年4月),包含预训练权重、对抗语料库(本工作称作ADC-2026)及全部可复现资源。
原文摘要 · Abstract (English)
State-of-the-art deepfake detectors achieve near-perfect in-domain accuracy yet degrade under cross-generator shifts, heavy compression, and adversarial perturbations. The core limitation remains the decoupling of semantic artifact learning from physical invariants: optical-flow discontinuities, specular-reflection inconsistencies, and cardiac-modulated reflectance (rPPG) are treated either as post-hoc features or ignored. We introduce PhyLAA-X, a novel physics-conditioned extension of Localized Artifact Attention (LAA-X). PhyLAA-X injects three end-to-end differentiable physics-derived feature volumes - optical-flow curl, specular-reflectance skewness, and spatially-upsampled rPPG power spectra - directly into the LAA-X attention computation via cross-attention gating and a resonance consistency loss. This forces the network to learn manipulation boundaries where semantic inconsistencies and physical violations co-occur - regions inherently harder for generative models to replicate consistently. PhyLAA-X is embedded across an efficient spatiotemporal ensemble (EfficientNet-B4+BiLSTM, ResNeXt-101+Transformer, Xception+causal Conv1D) with uncertainty-aware adaptive weighting. On FaceForensics++ (c23), Aletheia reaches 97.2% accuracy / 0.992 AUC-ROC; on Celeb-DF v2, 94.9% / 0.981; on DFDC, 90.8% / 0.966 - outperforming the strongest published baseline (LAA-Net [1]) by 4.1-7.3% in cross-generator settings and maintaining 79.4% accuracy under epsilon = 0.02 PGD-10 attacks. Single-backbone ablations confirm PhyLAA-X alone delivers a 4.2% cross-dataset AUC gain. The full production system is open-sourced at https://github.com/devghori1264/Aletheia (v1.2, April 2026) with pretrained weights, the adversarial corpus (referred to as ADC-2026 in this work), and complete reproducibility artifacts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。