提出轻量级多流声码器,实现实时语音合成低延迟高保真。
Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments
- 通过多流分解改进Wavehax,提升流式处理效率。
- 在仅用CPU环境下验证,实现低延迟与高吞吐的平衡。
- 模型小巧易部署,适合资源受限场景使用。
在实时语音合成中,神经声码器需通过因果处理和流式传输实现低延迟。然而,流式处理引入了批量合成中不存在的效率问题,如并行度受限、帧间依赖管理困难及参数加载开销。本文提出多流Wavehax(MS-Wavehax),通过将无混叠声码器Wavehax扩展为多流分解结构,构建高效低延迟流式声码器。我们在仅使用CPU的环境中分析了延迟-吞吐权衡关系,识别出流式神经声码器的关键瓶颈。研究结果为优化分块大小及设计适配特定应用需求与硬件限制的声码器提供了实用指导。此外,主观评估表明,MS-Wavehax在因果与非因果条件下均能提供高质量语音输出,同时具备极小模型体积,易于在资源受限环境中部署。
原文摘要 · Abstract (English)
In real-time speech synthesis, neural vocoders often require low-latency synthesis through causal processing and streaming. However, streaming introduces inefficiencies absent in batch synthesis, such as limited parallelism, inter-frame dependency management, and parameter loading overhead. This paper proposes multi-stream Wavehax (MS-Wavehax), an efficient neural vocoder for low-latency streaming, by extending the aliasing-free neural vocoder Wavehax with multi-stream decomposition. We analyze the latency-throughput trade-off in a CPU-only environment and identify key bottlenecks in streaming neural vocoders. Our findings provide practical insights for optimizing chunk sizes and designing vocoders tailored to specific application demands and hardware constraints. Furthermore, our subjective evaluations show that MS-Wavehax delivers high speech quality under causal and non-causal conditions while being remarkably compact and easily deployable in resource-constrained environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。