arXiv:2507.22781cs.CV2025-07被引 4

HOLA通过多模态预训练和分层上下文建模,提升视频级伪造检测精度。

HOLA: Enhancing Audio-visual Deepfake Detection via Hierarchical Contextual Aggregations and Efficient Pre-training

  • 采用自建181万样本数据集进行视听自监督预训练。
  • 在TestA上达0.9635 AUC,领先第二名0.0476。
  • 适合关注视频伪造检测与多模态模型的开发者。

生成式AI的发展使视频级深度伪造检测愈发困难,暴露了现有技术的局限性。本文提出HOLA,针对2025年1M-Deepfakes检测挑战赛的视频级伪造检测任务。受通用领域大规模预训练成功的启发,我们首次在多模态视频级伪造检测中实现大规模视听自监督预训练,利用自建181万样本数据集,构建统一的两阶段框架。HOLA包含迭代感知的跨模态学习模块,用于选择性视听交互;基于局部-全局视角的门控分层上下文建模;以及金字塔式精炼器,实现尺度感知的跨粒度语义增强。此外,提出伪监督信号注入策略进一步提升性能。大量实验表明,所提HOLA在专家模型与多模态大模型(MLLMs)上均表现优异。消融实验验证了各组件的关键作用。显著的是,HOLA在TestA集上取得0.9635 AUC,排名第一,超越第二名0.0476 AUC。

原文摘要 · Abstract (English)

Advances in Generative AI have made video-level deepfake detection increasingly challenging, exposing the limitations of current detection techniques. In this paper, we present HOLA, our solution to the Video-Level Deepfake Detection track of 2025 1M-Deepfakes Detection Challenge. Inspired by the success of large-scale pre-training in the general domain, we first scale audio-visual self-supervised pre-training in the multimodal video-level deepfake detection, which leverages our self-built dataset of 1.81M samples, thereby leading to a unified two-stage framework. To be specific, HOLA features an iterative-aware cross-modal learning module for selective audio-visual interactions, hierarchical contextual modeling with gated aggregations under the local-global perspective, and a pyramid-like refiner for scale-aware cross-grained semantic enhancements. Moreover, we propose the pseudo supervised singal injection strategy to further boost model performance. Extensive experiments across expert models and MLLMs impressivly demonstrate the effectiveness of our proposed HOLA. We also conduct a series of ablation studies to explore the crucial design factors of our introduced components. Remarkably, our HOLA ranks 1st, outperforming the second by 0.0476 AUC on the TestA set.

伪造检测多模态自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。