arXiv:2409.05799cs.SDcs.CL2024-09中稿 · SLT被引 4

解决语音识别中发音特征带来的偏差问题,提升说话人验证准确性

PDAF: A Phonetic Debiasing Attention Framework For Speaker Verification

  • 引入音素去偏注意力框架,动态调整音素权重
  • 在X-Vector系统上使SRE19测试集的EER降低0.8%
  • 适合关注语音验证鲁棒性的研究者与工程师

说话人验证系统通过语音认证身份至关重要。传统方法仅比较特征向量,忽视了语音内容的影响。本文指出音素主导性(即音素出现频率或持续时间)是说话人验证中的关键线索。为此提出一种新型音素去偏注意力框架(PDAF),可集成至现有注意力机制中,缓解因音素主导性带来的偏差。PDAF通过调整每个音素的权重影响特征提取过程,实现对语音更精细的分析。实验表明,该方法显著提升了验证性能;在X-Vector基础上结合PDAF后,SRE19测试集的等错误率(EER)下降0.8%。此外,通过多种加权策略评估了音素特征对系统性能的影响。

原文摘要 · Abstract (English)

Speaker verification systems are crucial for authenticating identity through voice. Traditionally, these systems focus on comparing feature vectors, overlooking the speech's content. However, this paper challenges this by highlighting the importance of phonetic dominance, a measure of the frequency or duration of phonemes, as a crucial cue in speaker verification. A novel Phoneme Debiasing Attention Framework (PDAF) is introduced, integrating with existing attention frameworks to mitigate biases caused by phonetic dominance. PDAF adjusts the weighting for each phoneme and influences feature extraction, allowing for a more nuanced analysis of speech. This approach paves the way for more accurate and reliable identity authentication through voice. Furthermore, by employing various weighting strategies, we evaluate the influence of phonetic features on the efficacy of the speaker verification system.

说话人验证注意力机制音素

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。