arXiv:2606.15117cs.MMcs.AI2026-06

用教师-学生框架提升音视频伪造检测模型跨域泛化能力

Teacher-Student Structure for Domain Adaptation in Ensemble Audio-Visual Video Deepfake Detection

论文配图:Teacher-Student Structure for Domain Adaptation in Ensemble Audio-Visual Video Deepfake Detection
图 1 · 摘自论文原文
  • 构建音视频联合检测模型,通过教师-学生结构实现跨域适应
  • 在三个未知数据集上分别提升4.09%、17.94%和0.5%的AUC表现
  • 仅用少量目标域数据即可适配,适合实际部署场景

生成式AI的快速发展带来了更逼真的音视频伪造内容,引发严重的隐私与社会问题。现有方法在单一域内表现良好,但在跨域场景下性能显著下降。为此,本文提出EAV-DFD模型,结合深度集成音视频特征与教师-学生框架的领域自适应机制,增强模型在未见域中的泛化能力。以FakeAVCeleb为源域,使用DFDC、Deepfake_TIMIT和PolyGlotFake作为目标域进行评估。实验表明,仅用少量目标域数据训练学生模型,即可在三组测试中分别提升4.09%、17.94%和0.5%的AUC值,有效实现跨域适应,并可识别被篡改的模态,具备真实应用场景潜力。

原文摘要 · Abstract (English)

The rapid advancement of generative AI models is leading to more realistic deepfake media, encompassing the manipulation of audio, video, or both. This raises severe privacy and societal concerns. Numerous studies in this area have yielded promising intra-domain results; however, these models frequently exhibit decreased efficacy when faced with data from dissimilar domains. Consequently, recent deepfake detection approaches focus on enhancing the generalization ability through multiple techniques that incorporate all input modalities, including audio, images, and their interactions. In this regard, we propose the EAV-DFD method, a generalized deep ensemble audio-visual model (EAV-DFD) combined with a domain adaptation mechanism utilizing a teacher-student framework to enhance the model's ability to perform and generalize effectively across unseen domains. To evaluate the model's performance, we used the FakeAVCeleb dataset as the primary domain and the DFDC, Deepfake_TIMIT, and PolyGlotFake datasets as an unseen domain. Our experimental results demonstrate that the proposed framework is efficient in domain adaptation, improving AUC performance of the model by 4.09%, 17.94%, and 0.5% on three unseen datasets, using only a small portion of them to train the student model. This leads to a novel deepfake detection model capable of adapting to new domains and interpreting which modality has been manipulated, highlighting the potential of our approach for real-world applications.

音视频伪造领域自适应教师-学生深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。