arXiv:2604.28022cs.CV2026-04

提出新检测标准,识别深层伪造中内容与源不一致的语义错位问题。

Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge

论文配图:Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge
图 1 · 摘自论文原文
  • 构建包含语义错位的新类别,模拟真实场景中的多模态伪造
  • 在FakeAVCeleb上测试发现现有模型对语义错位数据失效
  • 引入语义强化策略提升跨架构检测鲁棒性,适合安全与可信系统研究者

当前深度伪造检测多为二分类,难以反映音频、视频或两者同时被篡改的多样性。四类多模态检测虽能区分篡改类型,但存在新问题:模型可能仅依赖数据源完整性判断真伪,而忽略内容语义一致性。若伪造源于内容而非源文件,现有方法能否识别?本文提出新评估框架,在四类设置基础上引入新类别:真实音频-真实视频但存在语义错位(RARV-SMM)。基于FakeAVCeleb数据集评估表明,主流模型在语义错位场景下表现显著下降。进一步设计三种RARV-SMM变体,揭示不同架构对模态偏差的脆弱性。还提出一种融合语义错位类别与ImageBind嵌入的语义强化策略,验证其在FakeAVCeleb和LAV-DF上的有效性,推动更真实的深度伪造检测发展。

原文摘要 · Abstract (English)

Current DeepFake detection scenarios are mostly binary, yet data manipulation can vary across audio, video, or both, whose variability is not captured in binary settings. Four-class audio-visual formulations address this by discriminating manipulation type, but introduce an unresolved problem: models may rely solely on data source integrity to detect DeepFakes without evaluating their semantic consistency. If the DeepFake origin is not in the data source but in its content, can semantic mismatch be assessed by the state-of-the-art? This paper proposes a new evaluation setup, extending the four-class formulation by explicitly modeling semantic-level inconsistency between authentic modalities with the introduction of a new class: Real Audio-Real Video with Semantic Mismatch RARV-SMM. We assess the robustness of state-of-the-art models in this new realistic DeepFake setting, using the FakeAVCeleb dataset, highlighting the limitations of existing approaches when faced with semantic mismatch data. We further introduce three RARV-SMM variants that expose distinct architectural vulnerabilities as audio-visual divergence increases. We also propose a semantic reinforcement strategy that incorporates the semantic mismatch class and ImageBind embeddings to probe whether an explicit semantic coherence signal improves detection across architectures with different detection strategies, on FakeAVCeleb and LAV-DF, contributing toward more realistic DeepFake detectors. The source code available at https://github.com/sharayu-20/deepfake-semantic-mismatch.

深度伪造语义一致性多模态检测图像绑定

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。