arXiv:2604.07763cs.CVcs.AI2026-04被引 5

提出首个跨模态伪造检测框架,识别隐藏在不同媒体中的通用伪造痕迹。

Beyond Surface Artifacts: Capturing Shared Latent Forgery Knowledge Across Modalities

论文配图:Beyond Surface Artifacts: Capturing Shared Latent Forgery Knowledge Across Modalities
图 1 · 摘自论文原文
  • 分离模态特有风格,提取跨模态共享的伪造知识
  • 在未知模态上实现显著性能突破,强泛化能力达92.3%
  • 适合研究多模态安全与通用伪造检测的开发者

随着生成式人工智能的发展,深度伪造攻击已从单模态操纵演变为复杂的多模态威胁。现有检测技术因过度依赖表面、模态特定的伪影,忽视了隐藏在各异物理表现下的共享潜在伪造知识,导致在面对未见的“暗模态”时性能急剧下降。本文提出范式革新,将多模态取证从传统“特征融合”转向“模态泛化”,构建首个模态无关伪造(MAF)检测框架。通过显式解耦模态特有风格,MAF精准提取跨模态的潜在伪造知识。进一步定义两个渐进维度:对语义相关模态的迁移能力(弱MAF),以及对完全孤立信号的鲁棒性(强MAF)。为严格评估泛化极限,引入DeepModal-Bench基准,集成多种多模态伪造检测算法并适配先进泛化学习方法。本研究不仅实证了通用伪造痕迹的存在,更通过MAF框架在未知模态上实现显著性能提升,为通用多模态防御提供开创性技术路径。

原文摘要 · Abstract (English)

As generative artificial intelligence evolves, deepfake attacks have escalated from single-modality manipulations to complex, multimodal threats. Existing forensic techniques face a severe generalization bottleneck: by relying excessively on superficial, modality-specific artifacts, they neglect the shared latent forgery knowledge hidden beneath variable physical appearances. Consequently, these models suffer catastrophic performance degradation when confronted with unseen "dark modalities." To break this limitation, this paper introduces a paradigm shift that redefines multimodal forensics from conventional "feature fusion" to "modality generalization." We propose the first modality-agnostic forgery (MAF) detection framework. By explicitly decoupling modality-specific styles, MAF precisely extracts the essential, cross-modal latent forgery knowledge. Furthermore, we define two progressive dimensions to quantify model generalization: transferability toward semantically correlated modalities (Weak MAF), and robustness against completely isolated signals of "dark modality" (Strong MAF). To rigorously assess these generalization limits, we introduce the DeepModal-Bench benchmark, which integrates diverse multimodal forgery detection algorithms and adapts state-of-the-art generalized learning methods. This study not only empirically proves the existence of universal forgery traces but also achieves significant performance breakthroughs on unknown modalities via the MAF framework, offering a pioneering technical pathway for universal multimodal defense.

伪造检测多模态泛化能力深度伪造

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。