arXiv:2601.17530cs.CL2026-01Conference of the …被引 1

用大模型对比学习提升多模态假视频检测能力

Revealing the Truth with ConLLM for Detecting Multi-Modal Deepfakes

论文配图:Revealing the Truth with ConLLM for Detecting Multi-Modal Deepfakes
图 1 · 摘自论文原文
  • 分两阶段:先提取各模态特征,再通过对比学习对齐并用大模型推理
  • 音频假视频误报率降50%,视频识别准确率提升8%,音视频任务增9%
  • 适合关注假信息检测、多模态分析的研究者和安全从业者

深度伪造技术的快速发展对社会与政治稳定构成严重威胁,其生成的超真实合成媒体可操纵公众认知。现有检测方法存在两大瓶颈:(1)模态碎片化,导致跨多样化且对抗性深伪模态泛化能力差;(2)浅层跨模态推理,难以捕捉细粒度语义不一致。为此,我们提出ConLLM(基于大语言模型的对比学习),一种用于鲁棒多模态深度伪造检测的混合框架。ConLLM采用两阶段架构:第一阶段使用预训练模型(PTMs)提取模态特定嵌入;第二阶段通过对比学习对齐这些嵌入以缓解模态碎片化,并利用大模型推理细化嵌入,以捕捉语义不一致。ConLLM在音频、视频及音视频模态上均表现优异,音频深伪检测等错误率(EER)降低最高达50%,视频准确率提升最高达8%,音视频任务准确率提升约9%。消融实验表明,基于PTM的嵌入在各模态上贡献了9%-10%的稳定提升。

原文摘要 · Abstract (English)

The rapid rise of deepfake technology poses a severe threat to social and political stability by enabling hyper-realistic synthetic media capable of manipulating public perception. However, existing detection methods struggle with two core limitations: (1) modality fragmentation, which leads to poor generalization across diverse and adversarial deepfake modalities; and (2) shallow inter-modal reasoning, resulting in limited detection of fine-grained semantic inconsistencies. To address these, we propose ConLLM (Contrastive Learning with Large Language Models), a hybrid framework for robust multimodal deepfake detection. ConLLM employs a two-stage architecture: stage 1 uses Pre-Trained Models (PTMs) to extract modality-specific embeddings; stage 2 aligns these embeddings via contrastive learning to mitigate modality fragmentation, and refines them using LLM-based reasoning to address shallow inter-modal reasoning by capturing semantic inconsistencies. ConLLM demonstrates strong performance across audio, video, and audio-visual modalities. It reduces audio deepfake EER by up to 50%, improves video accuracy by up to 8%, and achieves approximately 9% accuracy gains in audio-visual tasks. Ablation studies confirm that PTM-based embeddings contribute 9%-10% consistent improvements across modalities.

深度伪造检测多模态大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。