arXiv:2608.23363cs.CVcs.AI2026-08中稿 · BMVC 2026

用多模态专家模型提升伪造视频检测泛化能力

DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-Experts

论文配图:DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-Experts
图 1 · 摘自论文原文
  • 融合音频视觉多模态特征,通过专家混合架构增强判别力
  • 在5个数据集上均超越现有方法,跨域检测性能显著
  • 适合需要高泛化能力的深伪检测场景

音视频深度伪造检测是当前研究热点,主要挑战在于构建能跨生成方法泛化的检测器。我们推测,通过预训练模型从音视频中提取多种高层线索,可缓解过拟合问题。为此,我们集成多种预训练模型,提取嘴部运动、人脸分割、面部表情、头部姿态、注视追踪、心率、语音情感和语音活动等特征。进一步采用混合专家(MoE)主干网络融合单模态与多模态线索进行伪造检测。我们在五个深度伪造检测基准(MAVOS-DD、AVLips、PolyGlotFake、BioDeepAV、FakeAVCeleb)上开展域内与跨域实验,结果表明DF-MoE在所有对比方法中表现最优。代码已开源。

原文摘要 · Abstract (English)

Audio-visual deepfake detection is an actively studied topic, where one of the main challenges is to develop detectors able to generalize across deepfake generation methods. We conjecture that overfitting can be mitigated by extracting multiple high-level cues from the available audio and visual modalities via pre-trained models. We therefore assemble a wide variety of pre-trained models to extract features that encode mouth movements, face parsing, facial expressions, head pose, gaze tracking, heart rate, audio emotion and speech activity. We further integrate both unimodal and multimodal cues via a Mixture-of-Experts (MoE) backbone to detect deepfakes. We perform in-domain and cross-domain experiments on five benchmarks for deepfake detection (MAVOS-DD, AVLips, PolyGlotFake, BioDeepAV, FakeAVCeleb) to compare our framework (DF-MoE) with state-of-the-art methods. Our results indicate that DF-MoE obtains superior deepfake detection results, surpassing all competing methods. We release our code at https://github.com/vladhondru25/DF-MoE.

深度伪造检测多模态MoE泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。