arXiv:2601.21925cs.SD2026-01被引 1

通过关注音频片段内部结构,提升局部伪造语音定位精度。

Localizing Speech Deepfakes Beyond Transitions via Segment-Aware Learning

  • 引入段内位置标签与跨段混合增强,让模型聚焦片段内部特征。
  • 在非边界区域定位准确率显著提升,对过渡伪影依赖降低。
  • 适用于真实场景中复杂伪造音频的检测,尤其适合语音安全研究者。

局部深度伪造语音的定位仍具挑战性,因修改内容细微且分散。现有方法多依赖帧级预测,部分近期工作通过关注真实与伪造音频间的过渡区提升性能。然而我们发现这些模型过度依赖边界伪影,忽视后续被篡改内容。为此提出段意识学习(SAL)框架,强调理解整个片段内部结构。SAL引入两项核心技术:段内位置标签,基于片段内相对位置提供细粒度帧级监督;跨段混合,一种生成多样片段模式的数据增强方法。在多个深度伪造定位数据集上的实验表明,SAL在域内与域外设置下均表现优异,尤其在非边界区域取得显著提升,同时减少对过渡伪影的依赖。代码已开源。

原文摘要 · Abstract (English)

Localizing partial deepfake audio, where only segments of speech are manipulated, remains challenging due to the subtle and scattered nature of these modifications. Existing approaches typically rely on frame-level predictions to identify spoofed segments, and some recent methods improve performance by concentrating on the transitions between real and fake audio. However, we observe that these models tend to over-rely on boundary artifacts while neglecting the manipulated content that follows. We argue that effective localization requires understanding the entire segments beyond just detecting transitions. Thus, we propose Segment-Aware Learning (SAL), a framework that encourages models to focus on the internal structure of segments. SAL introduces two core techniques: Segment Positional Labeling, which provides fine-grained frame supervision based on relative position within a segment; and Cross-Segment Mixing, a data augmentation method that generates diverse segment patterns. Experiments across multiple deepfake localization datasets show that SAL consistently achieves strong performance in both in-domain and out-of-domain settings, with notable gains in non-boundary regions and reduced reliance on transition artifacts. The code is available at https://github.com/SentryMao/SAL.

语音伪造深度伪造音频安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。