arXiv:2603.23960cs.CV2026-03被引 1

通过挖掘音视频内在一致性,提升深伪检测的鲁棒性与泛化能力。

Leave No Stone Unturned: Uncovering Holistic Audio-Visual Intrinsic Coherence for Deepfake Detection

  • 基于真实视频预训练学习音视频结构与跨模态一致性先验
  • 在跨数据集场景下实现9.39% AP与9.37% AUC的显著提升
  • 适用于检测新兴商业生成器产生的文本到视频和图像到视频伪造

生成式AI的快速发展催生了高度逼真的音视频深伪内容,严重威胁个人安全与社会信任。现有检测方法多依赖单模态痕迹或音视频差异,难以联合利用双源信息;且依赖生成器特有痕迹的方法在面对未知伪造时泛化能力差。本文提出HAVIC,一种基于音视频内在一致性的深伪检测框架。HAVIC首先在真实视频上预训练,学习模态内结构一致性、模态间微观与宏观一致性先验;再通过动态融合策略进行全局自适应特征聚合。此外,我们构建了HiFi-AVDF数据集,涵盖当前主流商业生成器产生的文本到视频和图像到视频伪造。大量实验表明,HAVIC在多个基准上显著优于现有方法,在最具挑战性的跨数据集场景中分别提升9.39% AP与9.37% AUC。代码与数据集已开源。

原文摘要 · Abstract (English)

The rapid progress of generative AI has enabled hyper-realistic audio-visual deepfakes, intensifying threats to personal security and social trust. Most existing deepfake detectors rely either on uni-modal artifacts or audio-visual discrepancies, failing to jointly leverage both sources of information. Moreover, detectors that rely on generator-specific artifacts tend to exhibit degraded generalization when confronted with unseen forgeries. We argue that robust and generalizable detection should be grounded in intrinsic audio-visual coherence within and across modalities. Accordingly, we propose HAVIC, a Holistic Audio-Visual Intrinsic Coherence-based deepfake detector. HAVIC first learns priors of modality-specific structural coherence, inter-modal micro- and macro-coherence by pre-training on authentic videos. Based on the learned priors, HAVIC further performs holistic adaptive aggregation to dynamically fuse audio-visual features for deepfake detection. Additionally, we introduce HiFi-AVDF, a high-fidelity audio-visual deepfake dataset featuring both text-to-video and image-to-video forgeries from state-of-the-art commercial generators. Extensive experiments across several benchmarks demonstrate that HAVIC significantly outperforms existing state-of-the-art methods, achieving improvements of 9.39% AP and 9.37% AUC on the most challenging cross-dataset scenario. Our code and dataset are available at https://github.com/tuffy-studio/HAVIC.

深伪检测音视频一致跨域泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。