arXiv:2507.16596cs.CV2025-07被引 10

仅用视频级标注定位伪造片段,提升弱监督下时间局部化精度。

A Multimodal Deviation Perceiving Framework for Weakly-Supervised Temporal Forgery Localization

  • 设计跨模态注意力机制,捕捉音视频间时序一致性偏差。
  • 提出可扩展的偏差感知损失,增强伪造段与真实段的差异性。
  • 在弱监督下达到接近全监督的效果,适合大规模检测场景。

现有深度伪造检测研究多将任务视为分类或时间局部化问题,往往受限于标注成本高、耗时长且难以扩展至大规模数据集。为此,本文提出一种弱监督时间伪造定位的多模态偏差感知框架(MDP),仅需视频级标注即可识别伪造片段的时间起止点。该框架引入新颖的多模态交互机制(MI),通过保留时序特性的跨模态注意力,在概率嵌入空间中度量视觉与音频模态的相关性,从而发现模态间偏差并构建用于时间定位的综合视频特征。此外,设计可扩展的偏差感知损失,旨在放大伪造样本相邻片段间的偏差,同时缩小真实样本的偏差。大量实验表明,所提方法在多个评估指标上表现优异,效果接近全监督方法。

原文摘要 · Abstract (English)

Current researches on Deepfake forensics often treat detection as a classification task or temporal forgery localization problem, which are usually restrictive, time-consuming, and challenging to scale for large datasets. To resolve these issues, we present a multimodal deviation perceiving framework for weakly-supervised temporal forgery localization (MDP), which aims to identify temporal partial forged segments using only video-level annotations. The MDP proposes a novel multimodal interaction mechanism (MI) and an extensible deviation perceiving loss to perceive multimodal deviation, which achieves the refined start and end timestamps localization of forged segments. Specifically, MI introduces a temporal property preserving cross-modal attention to measure the relevance between the visual and audio modalities in the probabilistic embedding space. It could identify the inter-modality deviation and construct comprehensive video features for temporal forgery localization. To explore further temporal deviation for weakly-supervised learning, an extensible deviation perceiving loss has been proposed, aiming at enlarging the deviation of adjacent segments of the forged samples and reducing that of genuine samples. Extensive experiments demonstrate the effectiveness of the proposed framework and achieve comparable results to fully-supervised approaches in several evaluation metrics.

伪造检测弱监督多模态时间定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。