EffVOC用20毫秒低延迟实现高质量语音波形重建,优于现有方法。
EffVOC: Low-Delay Efficient Speech Waveform Reconstruction from Spectral Representations Without Phase

- 基于高效低延迟声码器,支持幅度谱或梅尔系数输入
- 在宽频/全频语音上达到4.17/4.15的主观评分,接近真实语音
- 延迟仅20毫秒,适合实时语音合成场景
Griffin-Lim算法虽在幅度谱重建相位方面具有开创性,但需无限长算法延迟。其低延迟变体语音质量较差。近期神经网络方法虽提升语音质量,但仍存在中高延迟且通常仅针对特定输入表示优化。本文提出EffVOC,基于高效低延迟声码器,支持从幅度谱或梅尔系数输入合成宽频(WB)或全频(FB)语音。我们在统一框架下评估多种模型规模,对比最新方法。结果表明,所提低延迟(20毫秒,相比32毫秒及以上)高效方法在幅度谱和梅尔系数输入下分别取得4.17/4.15(WB)和4.14/4.11(FB)的主观评分,位居前列,非常接近真实语音质量。
原文摘要 · Abstract (English)
The Griffin-Lim algorithm has been a seminal contribution for phase reconstruction from amplitude spectrograms, however, requiring (infinitely) high algorithmic delay. Its low-delay variant suffers in speech quality. Recent (generative) neural network methods improve on speech quality still at medium to high algorithmic delay, but often they are complex and optimized only for one specific input representation. We build upon an efficient low-delay speech vocoder and propose EffVOC, which supports synthesis of wideband (WB) or fullband (FB) speech from either amplitude spectrum or Mel coefficient inputs. We evaluate both input representations across multiple model sizes in a unified framework and compare against state of the art. Results show that our proposed low-delay (20 ms vs. 32 ms or more) efficient approach marks a new SOTA by achieving top-ranked subjective MOS scores (WB: 4.17/4.15, FB: 4.14/4.11) for amplitude spectrum/Mel representations, very close to ground truth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。