语音不是平稳过程,用周期平稳模型更准
Harmonics to the Rescue: Why Voiced Speech is Not a Wss Process
- 用周期平稳模型替代传统平稳模型建模语音
- 能提升频谱密度估计和声源分离效果
- 适合做语音处理与声学系统识别的研究者
语音处理算法常依赖对底层过程的统计假设。尽管多年研究,语音最合适的统计模型仍存争议。语音通常被建模为宽平稳(WSS)过程,但对频域相关的信号而言,该假设本质错误,因WSS隐含频谱无相关性。本文证明,有声语音更应视为周期平稳(CS)过程。采用CS模型而非WSS模型处理固有频域相关过程,可改善交叉功率谱密度(PSD)估计、声源分离与波束成形性能。我们展示CS过程中谐波频率间的相关性如何增强系统辨识能力,并通过仿真与真实语音数据验证了结论。
原文摘要 · Abstract (English)
Speech processing algorithms often rely on statistical knowledge of the underlying process. Despite many years of research, however, the debate on the most appropriate statistical model for speech still continues. Speech is commonly modeled as a wide-sense stationary (WSS) process. However, the use of the WSS model for spectrally correlated processes is fundamentally wrong, as WSS implies spectral uncorrelation. In this paper, we demonstrate that voiced speech can be more accurately represented as a cyclostationary (CS) process. By employing the CS rather than the WSS model for processes that are inherently correlated across frequency, it is possible to improve the estimation of cross-power spectral densities (PSDs), source separation, and beamforming. We illustrate how the correlation between harmonic frequencies of CS processes can enhance system identification, and validate our findings using both simulated and real speech data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。