arXiv:2506.18691cs.SDeess.AS2025-06

按音素级别评估语音增强算法在男女声上的表现差异。

Evaluating Multichannel Speech Enhancement Algorithms at the Phoneme Scale Across Genders

  • 在音素层面分析男女声的语音增强效果差异
  • 女性语音在爆破音、摩擦音和元音上干扰减少更明显
  • 适合关注性别差异对语音处理影响的研究者

多通道语音增强算法对于提升嘈杂环境下语音的可懂度至关重要。这些算法通常在话语层面进行评估,但该方法忽略了不同音素类别及男女说话人之间存在的声学特征差异。本文研究了性别和语音内容对语音增强算法的影响,通过分析音素和性别特异的谱特征来支持这一研究思路。实验表明,虽然话语层面的性别差异较小,但在音素层面存在显著差异:所测试的算法在女性语音上表现出更强的干扰抑制能力,且在爆破音、摩擦音和元音上产生的伪影更少;同时,在感知质量和语音识别指标上,对女性语音的性能也更高。

原文摘要 · Abstract (English)

Multichannel speech enhancement algorithms are essential for improving the intelligibility of speech signals in noisy environments. These algorithms are usually evaluated at the utterance level, but this approach overlooks the disparities in acoustic characteristics that are observed in different phoneme categories and between male and female speakers. In this paper, we investigate the impact of gender and phonetic content on speech enhancement algorithms. We motivate this approach by outlining phoneme- and gender-specific spectral features. Our experiments reveal that while utterance-level differences between genders are minimal, significant variations emerge at the phoneme level. Results show that the tested algorithms better reduce interference with fewer artifacts on female speech, particularly in plosives, fricatives, and vowels. Additionally, they demonstrate greater performance for female speech in terms of perceptual and speech recognition metrics.

语音增强音素级评估性别差异

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。