语音增强效果受原始语音特性影响,共振峰幅度越高越易提升
Influence of Clean Speech Characteristics on Speech Enhancement Performance
- 提取音高、共振峰、响度等语音特征,分析其对增强难度的影响
- 共振峰幅度与增强增益正相关,幅度越高效果越好
- 同一说话人不同语句表现差异大,需考虑个体语音变异性
语音增强(SE)性能通常被认为受噪声特性和信噪比影响,但原始语音信号的内在特性仍研究不足。本文系统分析了多个先进语音增强模型、语言和噪声条件下,清洁语音特征对增强难度的影响。从清洁语音中提取音高、共振峰、响度和频谱变化率特征,并计算其与客观评估指标(包括频率加权分段信噪比和PESQ)的相关性。结果表明,共振峰幅度始终能有效预测增强性能,共振峰越高且越稳定,增强增益越大。此外,即使在同一说话人内部,不同语句间的性能差异也显著,凸显了说话人内声学变异的重要性。这些发现为语音增强挑战提供了新视角,建议在数据集设计、评估协议和模型构建中纳入语音内在特性。
原文摘要 · Abstract (English)
Speech enhancement (SE) performance is known to depend on noise characteristics and signal to noise ratio (SNR), yet intrinsic properties of the clean speech signal itself remain an underexplored factor. In this work, we systematically analyze how clean speech characteristics influence enhancement difficulty across multiple state of the art SE models, languages, and noise conditions. We extract a set of pitch, formant, loudness, and spectral flux features from clean speech and compute correlations with objective SE metrics, including frequency weighted segmental SNR and PESQ. Our results show that formant amplitudes are consistently predictive of SE performance, with higher and more stable formants leading to larger enhancement gains. We further demonstrate that performance varies substantially even within a single speaker's utterances, highlighting the importance of intraspeaker acoustic variability. These findings provide new insights into SE challenges, suggesting that intrinsic speech characteristics should be considered when designing datasets, evaluation protocols, and enhancement models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。