arXiv:2606.08669cs.SDcs.LG2026-06

对比多种特征提取与分类器组合,发现语音欺骗检测需考虑语言和领域差异。

A Comparison of SSL-Based Feature Extractors and Back-End Classifiers for Spoofing Detection: A Multi-Corpus Training and Cross-Linguistic Analysis

论文配图:A Comparison of SSL-Based Feature Extractors and Back-End Classifiers for Spoofing Detection: A Multi-Corpus Training and Cross-Linguistic Analysis
图 1 · 摘自论文原文
  • 用四种自监督特征提取器配四类分类器,跨语种多数据集测试
  • 发现ASVspoof 5数据集存在领域偏差,简单缩放反而降低性能
  • 仅用8小时目标语言数据微调,即可显著提升检测鲁棒性

语音生物识别系统面临日益增长的欺骗攻击威胁,但检测模型的评估在不同数据集间缺乏一致性。为探究这种不可预测的性能波动,我们对四种自监督学习特征提取器与四种后端分类器进行了全面基准测试。比较了ResNet的层次化局部特征提取与注意力机制及图结构后端的全局序列与关系建模能力。通过在三种场景下跨三个数据集进行多语料训练及六组评估,实证分析得出两项关键发现:其一,揭示了ASVspoof 5数据集中存在的领域偏差,表明简单的数据缩放会主动损害性能;其二,跨语言分析显示,仅使用8小时目标语言数据进行微调,即可显著增强检测鲁棒性。这些结果强调了在欺骗检测中实施领域感知与语言特异性适配的必要性。

原文摘要 · Abstract (English)

Voice biometric systems face growing threats from spoofing attacks, yet the evaluation of detection models remains inconsistent across datasets. To investigate these unpredictable fluctuations, we conduct a comprehensive benchmark of four self-supervised learning feature extractors paired with four back-end classifiers. We compare the hierarchical local feature extraction of ResNet with the global sequence and relational modeling of attention and graph-based back-ends. Through multi-corpus training across three scenarios and six evaluation datasets, our empirical analysis yields two critical findings. First, we expose a domain bias within the ASVspoof 5 dataset, showing that naive data scaling actively degrades performance. Second, our cross-linguistic analysis reveals that fine-tuning with just 8 hours of target-language data enhances detection robustness. Together, these findings emphasize the critical need for domain-aware and language-specific adaptation in spoofing detection.

语音安全自监督学习欺骗检测跨语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。