arXiv:2505.01880cs.SDcs.CV2025-05IJCAI被引 7

弱监督下定位音频伪造片段,用语言协同学习提升精度

Weakly-supervised Audio Temporal Forgery Localization via Progressive Audio-language Co-learning Network

  • 通过音频与语言协同学习,动态融合语义先验捕捉伪造特征
  • 在三个公开数据集上达到当前最优性能,无需精细标注
  • 适合研究音频安全、弱监督学习的学者和工程师

音频时间伪造定位(ATFL)旨在精准识别被人为篡改的音频片段。现有方法依赖于代价高昂且难以获取的细粒度标注进行训练。为此,本文提出一种渐进式音频-语言协同学习网络(LOCO),通过协同学习与自监督机制,在弱监督场景下提升定位性能。具体而言,设计音频-语言协同学习模块,从时序与全局视角对齐语义,生成基于话语级标注与可学习提示的伪造感知提示,动态融入时序内容特征。此外,引入伪造定位模块,基于融合的伪造类激活序列生成伪造候选区域。最后,采用渐进式优化策略生成伪帧级标签,并利用有监督语义对比学习强化真实与伪造内容间的语义差异,持续优化伪造感知特征。大量实验表明,所提方法在三个公开基准上均取得当前最优性能。

原文摘要 · Abstract (English)

Audio temporal forgery localization (ATFL) aims to find the precise forgery regions of the partial spoof audio that is purposefully modified. Existing ATFL methods rely on training efficient networks using fine-grained annotations, which are obtained costly and challenging in real-world scenarios. To meet this challenge, in this paper, we propose a progressive audio-language co-learning network (LOCO) that adopts co-learning and self-supervision manners to prompt localization performance under weak supervision scenarios. Specifically, an audio-language co-learning module is first designed to capture forgery consensus features by aligning semantics from temporal and global perspectives. In this module, forgery-aware prompts are constructed by using utterance-level annotations together with learnable prompts, which can incorporate semantic priors into temporal content features dynamically. In addition, a forgery localization module is applied to produce forgery proposals based on fused forgery-class activation sequences. Finally, a progressive refinement strategy is introduced to generate pseudo frame-level labels and leverage supervised semantic contrastive learning to amplify the semantic distinction between real and fake content, thereby continuously optimizing forgery-aware features. Extensive experiments show that the proposed LOCO achieves SOTA performance on three public benchmarks.

音频伪造弱监督协同学习定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。