arXiv:2602.23393cs.SDcs.CV2026-02

用大模型统一检测音视频伪造,效果超越现有方法。

Leveraging large multimodal models for audio-video deepfake detection: a pilot study

论文配图:Leveraging large multimodal models for audio-video deepfake detection: a pilot study
图 1 · 摘自论文原文
  • 基于大模型构建多模态检测框架,通过提示词完成真假判断
  • 在FakeAVCeleb和Mavos-DD上达到新最好性能,尤其在Mavos-DD领先
  • 适合关注大模型应用与跨域泛化能力的研究者

音视频深度伪造检测(AVD)日益重要,因现代生成技术可制造逼真视听内容。当前多数多模态检测器为小型专用模型,在精心筛选的数据集上表现良好,但扩展性差、跨领域泛化能力弱。本文提出AV-LMMDetect,一个基于Qwen 2.5 Omni的监督微调(SFT)大型多模态模型,将AVD任务转化为“该视频是真是假?”的提示式二分类。模型联合分析音频与视觉流,采用两阶段训练:轻量级LoRA对齐后,进行音视频编码器全量微调。在FakeAVCeleb和Mavos-DD数据集上,AV-LMMDetect表现匹配或超越先前方法,并在Mavos-DD上创下新纪录。

原文摘要 · Abstract (English)

Audio-visual deepfake detection (AVD) is increasingly important as modern generators can fabricate convincing speech and video. Most current multimodal detectors are small, task-specific models: they work well on curated tests but scale poorly and generalize weakly across domains. We introduce AV-LMMDetect, a supervised fine-tuned (SFT) large multimodal model that casts AVD as a prompted yes/no classification - "Is this video real or fake?". Built on Qwen 2.5 Omni, it jointly analyzes audio and visual streams for deepfake detection and is trained in two stages: lightweight LoRA alignment followed by audio-visual encoder full fine-tuning. On FakeAVCeleb and Mavos-DD, AV-LMMDetect matches or surpasses prior methods and sets a new state of the art on Mavos-DD datasets.

音视频伪造大模型多模态检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。