用词级分类替代时间回归,高效定位语音伪造段落。
Word-Anchored Temporal Forgery Localization
- 将伪造定位转为词级二分类,对齐语音语义边界。
- 在跨数据集测试中准确率显著超越现有方法。
- 轻量级设计适合实际部署,抗类别不平衡强。
当前时间伪造定位(TFL)方法多依赖时间边界回归或连续帧级异常检测,存在特征粒度不匹配和计算成本高的问题。本文提出词锚定时间伪造定位(WAFL),将TFL任务从时间回归与连续定位转变为离散词级二分类。通过分析伪造本质,识别出最小有意义单位——词标记,并对齐数据预处理与语音的自然语言边界。为适配强大的预训练基础模型,引入取证特征重对齐(FFR)模块,将预训练语义空间表示映射至判别性取证流形。后续轻量级线性分类器可高效完成二分类任务。针对伪造检测固有的极端类别不平衡问题,设计以伪影为中心的非对称损失(ACA),通过动态抑制大量真实样本梯度,非对称优先关注细微取证伪影,打破标准精度-召回权衡。大量实验表明,WAFL在跨数据集和同数据集设置下均显著优于现有方法,且参数量少、计算效率高。
原文摘要 · Abstract (English)
Current temporal forgery localization (TFL) approaches typically rely on temporal boundary regression or continuous frame-level anomaly detection paradigms to derive candidate forgery proposals. However, they suffer not only from feature granularity misalignment but also from costly computation. To address these issues, we propose word-anchored temporal forgery localization (WAFL), a novel paradigm that shifts the TFL task from temporal regression and continuous localization to discrete word-level binary classification. Specifically, we first analyze the essence of temporal forgeries and identify the minimum meaningful forgery units, word tokens, and then align data preprocessing with the natural linguistic boundaries of speech. To adapt powerful pre-trained foundation backbones for feature extraction, we introduce the forensic feature realignment (FFR) module, mapping representations from the pre-trained semantic space to a discriminative forensic manifold. This allows subsequent lightweight linear classifiers to efficiently perform binary classification and accomplish the TFL task. Furthermore, to overcome the extreme class imbalance inherent to forgery detection, we design the artifact-centric asymmetric (ACA) loss, which breaks the standard precision-recall trade-off by dynamically suppressing overwhelming authentic gradients while asymmetrically prioritizing subtle forensic artifacts. Extensive experiments demonstrate that WAFL significantly outperforms state-of-the-art approaches in localization performance under both in- and cross-dataset settings, while requiring substantially fewer learnable parameters and operating at high computational efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。