通过多尺度时序建模提升语音欺骗检测的鲁棒性
Robust Spoofed Speech Detection via Temporal Pyramid Modeling

- 设计时序金字塔结构,捕捉局部伪迹与全局语调异常
- 在PartialSpoof数据集上达到99.24% AUC、3.87% EER
- 适用于跨语言、跨域的语音伪造检测场景
语音欺骗检测面临真实合成、语音转换和重放攻击的严峻挑战,跨数据集泛化能力仍是主要瓶颈。本文提出时序金字塔适配器,通过并行设计不同感受野的时序卷积,捕获从局部伪迹到全局语调异常的多尺度欺骗特征。结合自监督XLS-R表征与前端适配器(梅尔、Sinc及时序金字塔),实现多尺度时序建模。在ASVspoof 2017、ASVspoof 2021(DF/LA)、PartialSpoof、DiffSSD及多语言HQ-MPSD等多个基准上评估,结果表明该模型在PartialSpoof数据集上取得99.24% AUC与3.87% EER,显著优于基线模型(如LCNN-BLSTM: 9.87% EER,TRACE: 8.08% EER)。多语言实验表明,欺骗特征与语言无关;尽管自监督表征提升鲁棒性,但在领域与语言迁移下性能下降,凸显需更优的适应与校准策略。
原文摘要 · Abstract (English)
Spoofed speech detection is increasingly challenged by realistic synthesis, voice conversion, and replay attacks, with cross-dataset generalization remaining a major limitation. This work we propose a Temporal Pyramid Adapter that utilize parallel temporal convolutions with varying receptive fields to capture multi-scale spoofing cues, ranging from local artifacts to global prosodic irregularities. We also integrated self-supervised XLS-R representations combined with front-end adapters, including Mel, Sinc, and a Temporal Pyramid design for multi-scale temporal modeling. The proposed model is evaluated cross multiple benchmark including ASVspoof 2017, ASVspoof 2021 (DF/LA), PartialSpoof, DiffSSD, and multilingual HQ-MPSD datasets. Experimental results demonstrate that Temporal Pyramid model obtained AUC of 99.24% and a EER of 3.87% on the PartialSpoof database, which is significantly outperforming the base model and several SOTA baseline such as LCNN-BLSTM (9.87% EER) and TRACE (8.08% EER). Additionally, multilingual evaluations confirm that while spoofing artifact are independent from language. While self-supervised representations improve robustness, performance degrades under domain and language shifts, highlighting the need for better adaptation and calibration strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。