arXiv:2603.25140cs.CVcs.AI2026-03

自监督学习检测音视频伪造,无需合成数据训练。

SAVe: Self-Supervised Audio-visual Deepfake Detection Exploiting Visual Artifacts and Audio-visual Misalignment

  • 自动生成保身份的伪伪造视频,学习多尺度视觉异常。
  • 通过音视频对齐分析,识别伪造特有的时序错位特征。
  • 仅用真实视频训练,泛化能力更强,适合新类型伪造检测。

多模态深度伪造可能产生细微视觉瑕疵和跨模态不一致,难以检测,尤其当检测器主要在精心构建的合成伪造数据上训练时。这类依赖会引入数据集与生成器偏差,限制模型可扩展性和对未知篡改的鲁棒性。我们提出SAVe,一种完全基于真实视频的自监督音视频伪造检测框架。SAVe通过即时生成保持身份、区域感知的伪混合伪造样本,模拟篡改痕迹,使模型学习多层级面部粒度的互补视觉线索。为捕捉跨模态证据,该框架还引入音视频对齐组件,检测唇语与语音之间的时序错位模式,这是音视频伪造的典型特征。在FakeAVCeleb和AV-LipSync-TIMIT数据集上的实验表明,SAVe在域内表现具有竞争力,并展现出强大的跨数据集泛化能力,凸显自监督学习在多模态伪造检测中的可扩展性优势。

原文摘要 · Abstract (English)

Multimodal deepfakes can exhibit subtle visual artifacts and cross-modal inconsistencies, which remain challenging to detect, especially when detectors are trained primarily on curated synthetic forgeries. Such synthetic dependence can introduce dataset and generator bias, limiting scalability and robustness to unseen manipulations. We propose SAVe, a self-supervised audio-visual deepfake detection framework that learns entirely on authentic videos. SAVe generates on-the-fly, identity-preserving, region-aware self-blended pseudo-manipulations to emulate tampering artifacts, enabling the model to learn complementary visual cues across multiple facial granularities. To capture cross-modal evidence, SAVe also models lip-speech synchronization via an audio-visual alignment component that detects temporal misalignment patterns characteristic of audio-visual forgeries. Experiments on FakeAVCeleb and AV-LipSync-TIMIT demonstrate competitive in-domain performance and strong cross-dataset generalization, highlighting self-supervised learning as a scalable paradigm for multimodal deepfake detection.

音视频伪造自监督学习多模态检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。