arXiv:2511.21237cs.CV2025-11

提出三层次框架,精准定位部分音频伪造的细微痕迹。

3-Tracer: A Tri-level Temporal-Aware Framework for Audio Forgery Detection and Localization

  • 分帧、段落、音频三层分析,捕捉不同时间尺度异常
  • 在三个数据集上达到当前最优检测性能
  • 适合音频安全、数字取证领域研究者使用

近期出现的部分音频伪造攻击中,攻击者仅修改语义关键帧,同时保持整体听觉真实性,使检测难度显著提升。现有方法多独立判断单帧真伪,缺乏跨时间层级的层次化分析能力。本文提出T3-Tracer框架,首次从帧、段落和音频三个层次联合分析音频,全面捕捉伪造痕迹。核心包含两个模块:帧-音频特征聚合模块(FA-FAM)融合帧内与全局时序信息,识别帧内伪造线索与语义不一致;段落级多尺度差异感知模块(SMDAM)采用双分支结构,在多尺度时间窗口内建模帧特征与帧间差异,有效定位伪造边界处的突变异常。在三个挑战性数据集上的实验表明,该方法性能达到当前最优。

原文摘要 · Abstract (English)

Recently, partial audio forgery has emerged as a new form of audio manipulation. Attackers selectively modify partial but semantically critical frames while preserving the overall perceptual authenticity, making such forgeries particularly difficult to detect. Existing methods focus on independently detecting whether a single frame is forged, lacking the hierarchical structure to capture both transient and sustained anomalies across different temporal levels. To address these limitations, We identify three key levels relevant to partial audio forgery detection and present T3-Tracer, the first framework that jointly analyzes audio at the frame, segment, and audio levels to comprehensively detect forgery traces. T3-Tracer consists of two complementary core modules: the Frame-Audio Feature Aggregation Module (FA-FAM) and the Segment-level Multi-Scale Discrepancy-Aware Module (SMDAM). FA-FAM is designed to detect the authenticity of each audio frame. It combines both frame-level and audio-level temporal information to detect intra-frame forgery cues and global semantic inconsistencies. To further refine and correct frame detection, we introduce SMDAM to detect forgery boundaries at the segment level. It adopts a dual-branch architecture that jointly models frame features and inter-frame differences across multi-scale temporal windows, effectively identifying abrupt anomalies that appeared on the forged boundaries. Extensive experiments conducted on three challenging datasets demonstrate that our approach achieves state-of-the-art performance.

音频伪造检测定位多尺度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。