arXiv:2509.17091cs.SDcs.CL2025-09EMNLP被引 8

构建首个全面评估语音识别鲁棒性的基准,覆盖真实场景下多种干扰因素。

SVeritas: Benchmark for Robust Speaker Verification under Diverse Conditions

  • 设计多维度压力测试框架,涵盖噪声、设备差异、年龄语言等30+真实场景
  • 发现主流模型在跨语言和编码压缩下性能下降超20个百分点
  • 揭示不同年龄、性别、语言群体间的识别差异,推动公平性改进

语音验证(SV)模型正广泛应用于安全、个性化和访问控制,但其对真实世界挑战的鲁棒性尚未得到充分评估。这些挑战包括自然与恶意导致的信号退化或注册与测试数据的不匹配。现有基准仅覆盖部分条件,遗漏关键问题。我们提出SVeritas,一个全面的语音验证评估基准套件,涵盖录音时长、即兴程度、内容、噪声、麦克风距离、混响、信道不匹配、音频带宽、编码器、说话人年龄,以及对欺骗和对抗攻击的敏感性。尽管已有多个基准分别覆盖部分问题,但SVeritas首次系统整合全部因素,并引入数个此前未被评测的重要真实场景。通过SVeritas评估多个先进SV模型,发现某些架构在常见失真下表现稳定,但在跨语言试验、年龄不匹配和编码压缩场景中性能显著下降。进一步按人口统计子群分析,揭示了年龄、性别和语言背景间的鲁棒性差异。通过标准化真实与合成压力条件下的评估,SVeritas可精准诊断模型弱点,为提升公平可靠语音验证系统奠定基础。

原文摘要 · Abstract (English)

Speaker verification (SV) models are increasingly integrated into security, personalization, and access control systems, yet their robustness to many real-world challenges remains inadequately benchmarked. These include a variety of natural and maliciously created conditions causing signal degradations or mismatches between enrollment and test data, impacting performance. Existing benchmarks evaluate only subsets of these conditions, missing others entirely. We introduce SVeritas, a comprehensive Speaker Verification tasks benchmark suite, assessing SV systems under stressors like recording duration, spontaneity, content, noise, microphone distance, reverberation, channel mismatches, audio bandwidth, codecs, speaker age, and susceptibility to spoofing and adversarial attacks. While several benchmarks do exist that each cover some of these issues, SVeritas is the first comprehensive evaluation that not only includes all of these, but also several other entirely new, but nonetheless important, real-life conditions that have not previously been benchmarked. We use SVeritas to evaluate several state-of-the-art SV models and observe that while some architectures maintain stability under common distortions, they suffer substantial performance degradation in scenarios involving cross-language trials, age mismatches, and codec-induced compression. Extending our analysis across demographic subgroups, we further identify disparities in robustness across age groups, gender, and linguistic backgrounds. By standardizing evaluation under realistic and synthetic stress conditions, SVeritas enables precise diagnosis of model weaknesses and establishes a foundation for advancing equitable and reliable speaker verification systems.

语音验证鲁棒性基准测试公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。