arXiv:2512.07352cs.SD2025-12中稿 · Interspeech 2026被引 3

构建多接口语音伪造检测数据集,提升模型在真实场景下的抗欺骗能力

MultiAPI Spoof: A Multi-API Dataset and Local-Attention Network for Speech Anti-spoofing Detection

  • 基于30个不同API生成230小时合成语音,模拟真实商用系统多样性
  • 提出局部注意力网络,显著提升对细微伪造痕迹的识别能力
  • 支持伪造音频来源溯源,适合语音安全与反欺诈研究者使用

现有语音反伪造基准依赖有限的公开模型,难以反映真实场景中商业系统使用的多样且常为专有的API。为此,我们提出MultiAPI Spoof,一个包含约230小时合成语音的多接口语音反伪造数据集,由30种不同API生成,涵盖商业服务、开源模型和在线平台。同时,我们提出Nes2Net-LA,一种增强局部注意力机制的Nes2Net变体,提升局部上下文建模与细粒度伪造特征提取能力。基于该数据集,我们定义了API溯源任务,实现对伪造音频生成源的细粒度定位。实验表明,Nes2Net-LA达到当前最佳性能,尤其在多样且未见过的伪造条件下表现出更强鲁棒性。代码与数据集已公开。

原文摘要 · Abstract (English)

Existing speech anti-spoofing benchmarks rely on a narrow set of public models, creating a substantial gap from real-world scenarios in which commercial systems employ diverse, often proprietary APIs. To address this issue, we introduce MultiAPI Spoof, a multi-API audio anti-spoofing dataset comprising about 230 hours of synthetic speech generated by 30 distinct APIs, including commercial services, open-source models, and online platforms. Furthermore, we propose Nes2Net-LA, a local-attention enhanced variant of Nes2Net that improves local context modeling and fine-grained spoofing feature extraction. Based on this dataset, we also define the API tracing task, enabling fine-grained attribution of spoofed audio to its generation source. Experiments show that Nes2Net-LA achieves state-of-the-art performance and offers superior robustness, particularly under diverse and unseen spoofing conditions. Code \footnote{https://github.com/XuepingZhang/MultiAPI-Spoof} and dataset \footnote{https://xuepingzhang.github.io/MultiAPI-Spoof-Dataset/} have been released.

语音安全反伪造多源检测注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。