用自监督音频+手工视觉特征,高效检测音视频伪造。
KLASSify to Verify: Audio-Visual Deepfake Detection Using SSL-based Audio and Handcrafted Visual Features
- 音频用自监督模型+图注意力网络,捕捉深层语义特征。
- 纯音频检测达92.78%准确率,时序定位IoU达0.3536。
- 兼顾性能与可解释性,适合实际部署场景。
语音驱动的虚拟人生成和先进文本转语音(TTS)模型的快速发展,催生了更复杂的时序深度伪造。现有顶尖检测方法虽准确,但计算开销大,难以泛化到新攻击手法。为此,我们提出针对AV-Deepfake1M 2025挑战的多模态方案:视觉模态采用手工特征以增强可解释性与适应性;音频模态则采用自监督学习(SSL)骨干网络结合图注意力机制,捕获丰富音频表征,提升检测鲁棒性。该方法在保持高性能的同时兼顾实际部署可行性,强调抗干扰能力与潜在可解释性。在AV-Deepfake1M++数据集上,仅使用音频模态即实现深度伪造分类任务92.78%的AUC,时序定位任务0.3536的IoU。
原文摘要 · Abstract (English)
The rapid development of audio-driven talking head generators and advanced Text-To-Speech (TTS) models has led to more sophisticated temporal deepfakes. These advances highlight the need for robust methods capable of detecting and localizing deepfakes, even under novel, unseen attack scenarios. Current state-of-the-art deepfake detectors, while accurate, are often computationally expensive and struggle to generalize to novel manipulation techniques. To address these challenges, we propose multimodal approaches for the AV-Deepfake1M 2025 challenge. For the visual modality, we leverage handcrafted features to improve interpretability and adaptability. For the audio modality, we adapt a self-supervised learning (SSL) backbone coupled with graph attention networks to capture rich audio representations, improving detection robustness. Our approach strikes a balance between performance and real-world deployment, focusing on resilience and potential interpretability. On the AV-Deepfake1M++ dataset, our multimodal system achieves AUC of 92.78% for deepfake classification task and IoU of 0.3536 for temporal localization using only the audio modality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。