分析语音频段抗噪能力,发现中频段最易受噪声影响
From the perspective of perceptual speech quality: The robustness of frequency bands to noise
- 用MUSHRA方法测试32个频段在不同信噪比下的感知质量
- 中频段在噪声下感知质量下降最明显,鲁棒性最差
- 研究结果对语音增强和通信系统优化有指导意义
语音质量是语音研究的重要方向,常与语音可懂度并列考察。尽管频段级感知可懂度已有较多研究,但语音质量的频段级分析仍不充分。本文提出一种受MUSHRA启发的方法,以感知语音质量为指标,评估各频率带在噪声下的鲁棒性。将语音信号分割为32个频段,并施加真实世界噪声,分别在不同信噪比下进行测试。基于人工评分的感知质量得分,计算各频段的抗噪指数。结果显示,中频区域在噪声下感知质量退化最显著,鲁棒性最弱。该发现提示未来提升语音质量的研究应重点关注中频段。
原文摘要 · Abstract (English)
Speech quality is one of the main foci of speech-related research, where it is frequently studied with speech intelligibility, another essential measurement. Band-level perceptual speech intelligibility, however, has been studied frequently, whereas speech quality has not been thoroughly analyzed. In this paper, a Multiple Stimuli With Hidden Reference and Anchor (MUSHRA) inspired approach was proposed to study the individual robustness of frequency bands to noise with perceptual speech quality as the measure. Speech signals were filtered into thirty-two frequency bands with compromising real-world noise employed at different signal-to-noise ratios. Robustness to noise indices of individual frequency bands was calculated based on the human-rated perceptual quality scores assigned to the reconstructed noisy speech signals. Trends in the results suggest the mid-frequency region appeared less robust to noise in terms of perceptual speech quality. These findings suggest future research aiming at improving speech quality should pay more attention to the mid-frequency region of the speech signals accordingly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。