利用生成模型内部音视频对齐信号,提升深度伪造检测鲁棒性。
X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake Detection
- 通过DDIM反演获取生成模型内部的音视频注意力特征
- 在新数据集MMDF上准确率比现有方法高13.1%
- 适用于GAN、扩散模型等多种生成器,泛化能力强
当前生成系统制作的高保真合成视频激增,极大增加了滥用风险,挑战了人类与现有检测器。基于此,我们从生成器视角出发,发现其内部交叉注意力机制编码了精细的语音-动作对齐信息,可作为伪造检测的有用线索。为此,我们提出X-AVDT,一种鲁棒且可泛化的深度伪造检测器,通过DDIM反演访问生成器内部的音视频信号以揭示这些线索。X-AVDT提取两种互补信号:(i) 反演引入的不一致视频复合图像;(ii) 反映生成过程中模态对齐的音视频交叉注意力特征。为实现跨生成器的可信评估,我们进一步构建了MMDF,一个涵盖多种篡改类型和快速演进生成范式的多模态深度伪造数据集,包括GAN、扩散模型和流匹配。大量实验表明,X-AVDT在MMDF上表现领先,并在外部基准和未见生成器上展现强泛化能力,准确率提升13.1%。研究强调了利用内部音视频一致性线索对提升未来生成器检测鲁棒性的关键作用。
原文摘要 · Abstract (English)
The surge of highly realistic synthetic videos produced by contemporary generative systems has significantly increased the risk of malicious use, challenging both humans and existing detectors. Against this backdrop, we take a generator-side view and observe that internal cross-attention mechanisms in these models encode fine-grained speech-motion alignment, offering useful correspondence cues for forgery detection. Building on this insight, we propose X-AVDT, a robust and generalizable deepfake detector that probes generator-internal audio-visual signals accessed via DDIM inversion to expose these cues. X-AVDT extracts two complementary signals: (i) a video composite capturing inversion-induced discrepancies, and (ii) an audio-visual cross-attention feature reflecting modality alignment enforced during generation. To enable faithful cross-generator evaluation, we further introduce MMDF, a new multimodal deepfake dataset spanning diverse manipulation types and rapidly evolving synthesis paradigms, including GANs, diffusion, and flow-matching. Extensive experiments demonstrate that X-AVDT achieves leading performance on MMDF and generalizes strongly to external benchmarks and unseen generators, outperforming existing methods with accuracy improved by 13.1%. Our findings highlight the importance of leveraging internal audio-visual consistency cues for robustness to future generators in deepfake detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。