研究语音预测最大增益,发现声带振动语音可提升6dB
Analysis of the Maximum Prediction Gain of Short-Term Prediction on Sustained Speech
- 用核回归与信息论方法分析语音预测上限
- 非浊音语音线性预测已达最优,浊音语音可多增2-6dB
- 结果差异大,适合语音编码与信号处理研究者
信号预测广泛应用于经济预测、回声消除和数据压缩,尤其在语音与音乐的预测编码中。预测编码通过预测降低传输或存储所需的比特率。预测增益是信号编码中的经典指标,关联均方预测误差与预测编码器的信噪比。为评估预测模型,了解不依赖具体模型的最大预测增益至关重要。本文利用Nadaraya-Watson核回归(NWKR)与信息论上界,分析新采集的持续语音/音素数据集上的预测增益上限。结果显示,对于非浊音语音,线性预测器在最多0.3 dB内达到最大预测增益;而对于浊音语音,单抽头预测器最优仍为线性,但从两抽头起,最大可实现预测增益比线性预测高出约2至6 dB。不同说话人之间存在显著差异。所创建的数据集及代码可应要求用于研究。
原文摘要 · Abstract (English)
Signal prediction is widely used in, e.g., economic forecasting, echo cancellation and in data compression, particularly in predictive coding of speech and music. Predictive coding algorithms reduce the bit-rate required for data transmission or storage by signal prediction. The prediction gain is a classic measure in applied signal coding of the quality of a predictor, as it links the mean-squared prediction error to the signal-to-quantization-noise of predictive coders. To evaluate predictor models, knowledge about the maximum achievable prediction gain independent of a predictor model is desirable. In this manuscript, Nadaraya-Watson kernel-regression (NWKR) and an information theoretic upper bound are applied to analyze the upper bound of the prediction gain on a newly recorded dataset of sustained speech/phonemes. It was found that for unvoiced speech a linear predictor always achieves the maximum prediction gain within at most 0.3 dB. On voiced speech, the optimum one-tap predictor was found to be linear but starting with two taps, the maximum achievable prediction gain was found to be about 2 dB to 6 dB above the prediction gain of the linear predictor. Significant differences between speakers/subjects were observed. The created dataset as well as the code can be obtained for research purpose upon request.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。