arXiv:2603.01482eess.AScs.AI2026-03中稿 · ICASSP被引 4

构建首个自监督语音模型音频伪造检测基准,验证大模型在真实场景中的防御能力。

A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection

  • 构建Spoof-SUPERB基准,评估20种自监督语音模型在伪造检测上的表现。
  • 大尺度判别模型如XLS-R、WavLM Large在多数据集上显著领先,准确率超90%。
  • 适合安全研究者和语音系统开发者,为防伪造提供可复现的技术参考。

自监督学习(SSL)已重塑语音处理领域,如SUPERB基准使不同下游任务间具备公平比较基础。然而,音频伪造检测这一关键安全问题长期未被纳入此类评估体系。本文提出Spoof-SUPERB,首个面向音频伪造检测的自监督语音模型基准,系统评估20种涵盖生成型、判别型及频谱型架构的SSL模型。在多个域内与域外数据集上进行测试,结果表明:大规模判别型模型(如XLS-R、UniSpeech-SAT、WavLM Large)凭借多语言预训练、说话人感知目标和模型规模优势,持续领先,部分任务准确率达93.7%以上。进一步分析显示,生成型方法在声学退化下性能骤降,而判别型模型仍具强鲁棒性。该基准建立可复现基线,为提升语音系统安全性提供实证依据。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) has transformed speech processing, with benchmarks such as SUPERB establishing fair comparisons across diverse downstream tasks. Despite it's security-critical importance, Audio deepfake detection has remained outside these efforts. In this work, we introduce Spoof-SUPERB, a benchmark for audio deepfake detection that systematically evaluates 20 SSL models spanning generative, discriminative, and spectrogram-based architectures. We evaluated these models on multiple in-domain and out-of-domain datasets. Our results reveal that large-scale discriminative models such as XLS-R, UniSpeech-SAT, and WavLM Large consistently outperform other models, benefiting from multilingual pretraining, speaker-aware objectives, and model scale. We further analyze the robustness of these models under acoustic degradations, showing that generative approaches degrade sharply, while discriminative models remain resilient. This benchmark establishes a reproducible baseline and provides practical insights into which SSL representations are most reliable for securing speech systems against audio deepfakes.

音频伪造自监督学习语音安全模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。