arXiv:2604.20229cs.SDcs.AI2026-04

提升语音验证对耳语的鲁棒性,让隐私保护说话也能精准识别。

Enhancing Speaker Verification with Whispered Speech via Post-Processing

论文配图:Enhancing Speaker Verification with Whispered Speech via Post-Processing
图 1 · 摘自论文原文
  • 用编码器-解码器结构微调声纹模型,联合优化相似度与三元组损失。
  • 在正常与耳语对比测试中,错误率降低22.26%,AUC达98.16%。
  • 耳语对耳语验证误差仅1.88%,适合高安全场景或病理语音应用。

语音验证通过分析声音确认身份。耳语在声学特征上不同于正常发音,导致真实场景下系统性能下降,如为保护隐私、避免打扰他人或因疾病无法充分发声时。本文提出一种基于微调声纹主干网络的编码器-解码器模型,通过联合优化余弦相似度分类与三元组损失,增强对耳语干扰的鲁棒性。在正常语音与耳语对比测试中,相对基线(基线错误率6.77%)提升22.26%(我们5.27%),达到AUC 98.16%。在耳语对耳语测试中,错误率降至1.88%,AUC达99.73%,相较此前领先模型ReDimNet-B2提升15%。此外,我们总结了主流声纹模型在耳语下的表现,并评估噪声环境下性能:相同噪声水平对耳语语音的降维作用显著大于正常语音。

原文摘要 · Abstract (English)

Speaker verification is a task of confirming an individual's identity through the analysis of their voice. Whispered speech differs from phonated speech in acoustic characteristics, which degrades the performance of speaker verification systems in real-life scenarios, including avoiding fully phonated speech to protect privacy, disrupt others, or when the lack of full vocalization is dictated by a disease. In this paper we propose a model with a training recipe to obtain more robust representations against whispered speech hindrances. The proposed system employs an encoder--decoder structure built atop a fine-tuned speaker verification backbone, optimized jointly using cosine similarity--based classification and triplet loss. We gain relative improvement of 22.26\% compared to the baseline (baseline 6.77\% vs ours 5.27\%) in normal vs whispered speech trials, achieving AUC of 98.16\%. In tests comparing whispered to whispered, our model attains an EER of 1.88\% with AUC equal to 99.73\%, which represents a 15\% relative enhancement over the prior leading ReDimNet-B2. We also offer a summary of the most popular and state-of-the-art speaker verification models in terms of their performance with whispered speech. Additionally, we evaluate how these models perform under noisy audios, obtaining that generally the same relative level of noise degrades the performance of speaker verification more significantly on whispered speech than on normal speech.

声纹验证耳语识别语音鲁棒性后处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。