arXiv:2505.15233cs.CV2025-05被引 14

提出跨模态对齐与蒸馏框架,提升视频伪造检测准确率

CAD: A General Multimodal Framework for Video Deepfake Detection via Cross-Modal Alignment and Distillation

  • 通过跨模态对齐捕捉唇语与语音的语义不一致
  • 利用跨模态蒸馏保留各模态的细微伪造痕迹,融合更均衡
  • 在多模态和单模态数据集上均超越现有方法,适合伪造检测研究者

多模态深度伪造(视觉与听觉内容协同篡改)的兴起削弱了仅依赖单一模态痕迹或跨模态不一致性的现有检测器可靠性。本文首次证明,模态特异性伪造痕迹(如人脸替换伪影、频谱失真)与模态共享的语义错位(如唇语-语音不同步)提供互补证据,忽略任一都会限制检测性能。现有方法或简单融合模态特征而未调和其冲突,或过度关注语义错位而忽视细粒度伪造线索。为此,我们提出跨模态对齐与蒸馏(CAD)通用框架:1)跨模态对齐识别高层语义同步性不一致(如唇语-语音错配);2)跨模态蒸馏在融合时缓解特征冲突,同时保留模态特异性伪造痕迹(如合成音频中的频谱失真)。在多模态与单模态(如仅图像/仅视频)深度伪造基准上的大量实验表明,CAD显著优于先前方法,验证了多模态互补信息和谐整合的必要性。

原文摘要 · Abstract (English)

The rapid emergence of multimodal deepfakes (visual and auditory content are manipulated in concert) undermines the reliability of existing detectors that rely solely on modality-specific artifacts or cross-modal inconsistencies. In this work, we first demonstrate that modality-specific forensic traces (e.g., face-swap artifacts or spectral distortions) and modality-shared semantic misalignments (e.g., lip-speech asynchrony) offer complementary evidence, and that neglecting either aspect limits detection performance. Existing approaches either naively fuse modality-specific features without reconciling their conflicting characteristics or focus predominantly on semantic misalignment at the expense of modality-specific fine-grained artifact cues. To address these shortcomings, we propose a general multimodal framework for video deepfake detection via Cross-Modal Alignment and Distillation (CAD). CAD comprises two core components: 1) Cross-modal alignment that identifies inconsistencies in high-level semantic synchronization (e.g., lip-speech mismatches); 2) Cross-modal distillation that mitigates feature conflicts during fusion while preserving modality-specific forensic traces (e.g., spectral distortions in synthetic audio). Extensive experiments on both multimodal and unimodal (e.g., image-only/video-only)deepfake benchmarks demonstrate that CAD significantly outperforms previous methods, validating the necessity of harmonious integration of multimodal complementary information.

视频伪造多模态深度伪造检测跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。