arXiv:2506.06675eess.AScs.SD2025-06

分析语音中基频脉冲的幅相结构,评估三种轻量化声门重建方法。

Accurate analysis of the pitch pulse-based magnitude/phase structure of natural vowels and assessment of three lightweight time/frequency voicing restoration methods

  • 提出新算法分割基频脉冲,区分持续与连读元音特征。
  • 三类重建方法在听觉测试中表现差异显著,生理启发法最优。
  • 适合低资源设备实时语音修复,适用于语音康复与便携应用。

耳语语音因声带不振动而缺乏周期性信号成分,与自然语音的关键区别在于此。本文解决两个核心挑战:一是刻画并建模自然语音中连续基频脉冲的幅相结构演化,提出一种新型脉冲分割算法,揭示了持续与连读元音间的显著差异,为合成发声提供依据;二是基于模型的合成发声实现,比较三种方法——频域、时频联合、生理启发式独立滤波声门激励脉冲——通过客观指标与听觉测试验证其效果。结果表明,生理启发方法在词语语境下合成的持续与连读元音更具自然性,具备在低资源设备上实时、在线实现语音恢复的潜力。

原文摘要 · Abstract (English)

Whispered speech is produced when the vocal folds are not used, either intentionally, or due to a temporary or permanent voice condition. The essential difference between natural speech and whispered speech is that periodic signal components that exist in certain regions of the former, called voiced regions, as a consequence of the vibration of the vocal folds, are missing in the latter. The restoration of natural speech from whispered speech requires delicate signal processing procedures that are especially useful if they can be implemented on low-resourced portable devices, in real-time, and on-the-fly, taking advantage of the established source-filter paradigm of voice production and related models. This paper addresses two challenges that are intertwined and are key in informing and making viable this envisioned technological realization. The first challenge involves characterizing and modeling the evolution of the harmonic phase/magnitude structure of a sequence of individual pitch periods in a voiced region of natural speech comprising sustained or co-articulated vowels. This paper proposes a novel algorithm segmenting individual pitch pulses, which is then used to obtain illustrative results highlighting important differences between sustained and co-articulated vowels, and suggesting practical synthetic voicing approaches. The second challenge involves model-based synthetic voicing. Three implementation alternatives are described that differ in their signal reconstruction approaches: frequency-domain, combined frequency and time-domain, and physiologically-inspired separate filtering of glottal excitation pulses individually generated. The three alternatives are compared objectively using illustrative examples, and subjectively using the results of listening tests involving synthetic voicing of sustained and co-articulated vowels in word context.

语音重建轻量化声门建模耳语还原

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。