arXiv:2508.02179cs.CV2025-08被引 4

用视频级标注实现多模态细粒度伪造定位,提升弱监督检测精度。

Weakly Supervised Multimodal Temporal Forgery Localization via Multitask Learning

  • 多任务学习融合视觉与音频分类,仅需视频级标签。
  • 引入专家混合结构,提升定位精度与模型灵活性。
  • 设计时序感知损失函数,增强伪造片段的时序区分能力。

深度伪造视频的传播引发了信任危机并影响社会稳定。尽管已有众多方法用于深度伪造检测与定位,但针对弱监督多模态细粒度时间伪造定位(WS-MTFL)的系统性研究仍不足。本文提出一种基于多任务学习的弱监督多模态时间伪造定位方法(WMMT),在仅使用视频级标注的前提下,实现多模态细粒度深度伪造检测与时间局部化。将视觉与音频模态检测建模为两个二分类任务,通过多任务学习框架整合为多模态任务。WMMT采用专家混合结构,自适应选择特征与定位头,提升灵活性与定位精度;设计具有时序特性保持注意力机制的特征增强模块,识别跨模态与模态内特征偏差,构建综合视频特征。为进一步挖掘弱监督下的时序信息,提出可扩展的偏差感知损失,旨在放大伪造样本相邻片段间的差异,缩小真实样本的差异。大量实验表明,多任务学习在WS-MTFL中有效,WMMT在多个评估指标上达到与全监督方法相当的性能。

原文摘要 · Abstract (English)

The spread of Deepfake videos has caused a trust crisis and impaired social stability. Although numerous approaches have been proposed to address the challenges of Deepfake detection and localization, there is still a lack of systematic research on the weakly supervised multimodal fine-grained temporal forgery localization (WS-MTFL). In this paper, we propose a novel weakly supervised multimodal temporal forgery localization via multitask learning (WMMT), which addresses the WS-MTFL under the multitask learning paradigm. WMMT achieves multimodal fine-grained Deepfake detection and temporal partial forgery localization using merely video-level annotations. Specifically, visual and audio modality detection are formulated as two binary classification tasks. The multitask learning paradigm is introduced to integrate these tasks into a multimodal task. Furthermore, WMMT utilizes a Mixture-of-Experts structure to adaptively select appropriate features and localization head, achieving excellent flexibility and localization precision in WS-MTFL. A feature enhancement module with temporal property preserving attention mechanism is proposed to identify the intra- and inter-modality feature deviation and construct comprehensive video features. To further explore the temporal information for weakly supervised learning, an extensible deviation perceiving loss has been proposed, which aims to enlarge the deviation of adjacent segments of the forged samples and reduce the deviation of genuine samples. Extensive experiments demonstrate the effectiveness of multitask learning for WS-MTFL, and the WMMT achieves comparable results to fully supervised approaches in several evaluation metrics.

深度伪造多模态弱监督时序定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。