用单模态生成多模态特征,提升跨压缩率伪造检测鲁棒性
UMCL: Unimodal-generated Multimodal Contrastive Learning for Cross-compression-rate Deepfake Detection

- 从单一视觉数据生成三类互补特征,增强抗压缩能力
- 在多种压缩率下检测准确率超越现有方法,最低仍保持90%以上
- 适合需要高鲁棒性和可解释性的深度伪造检测应用场景
在深度伪造检测中,社交媒体平台采用的不同压缩程度严重挑战模型的泛化与可靠性。现有方法虽已从单模态发展到多模态,但仍存在局限:单模态方法在流媒体压缩下特征退化明显,而多模态方法依赖昂贵的数据采集与标注,且实际场景中各模态质量不一致或不可用。为此,我们提出一种新型无监督生成多模态对比学习框架(UMCL),用于鲁棒的跨压缩率(CCR)深度伪造检测。训练阶段,该方法将单一视觉模态转化为三类互补特征:抗压缩的rPPG信号、时序关键点动态和预训练视觉-语言模型的语义嵌入。通过亲和力驱动的语义对齐(ASA)策略,利用亲和矩阵建模跨模态关系,并通过对比学习优化其一致性。随后,跨质量相似性学习(CQSL)进一步提升特征在不同压缩率下的鲁棒性。大量实验表明,该方法在多种压缩率和篡改类型下均表现卓越,建立了新基准。尤其在个别特征退化时仍能保持高检测精度,且通过显式对齐提供可解释的特征关系分析。
原文摘要 · Abstract (English)
In deepfake detection, the varying degrees of compression employed by social media platforms pose significant challenges for model generalization and reliability. Although existing methods have progressed from single-modal to multimodal approaches, they face critical limitations: single-modal methods struggle with feature degradation under data compression in social media streaming, while multimodal approaches require expensive data collection and labeling and suffer from inconsistent modal quality or accessibility in real-world scenarios. To address these challenges, we propose a novel Unimodal-generated Multimodal Contrastive Learning (UMCL) framework for robust cross-compression-rate (CCR) deepfake detection. In the training stage, our approach transforms a single visual modality into three complementary features: compression-robust rPPG signals, temporal landmark dynamics, and semantic embeddings from pre-trained vision-language models. These features are explicitly aligned through an affinity-driven semantic alignment (ASA) strategy, which models inter-modal relationships through affinity matrices and optimizes their consistency through contrastive learning. Subsequently, our cross-quality similarity learning (CQSL) strategy enhances feature robustness across compression rates. Extensive experiments demonstrate that our method achieves superior performance across various compression rates and manipulation types, establishing a new benchmark for robust deepfake detection. Notably, our approach maintains high detection accuracy even when individual features degrade, while providing interpretable insights into feature relationships through explicit alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。