用低帧率高维连续符表示语音,实现稳定高保真生成
Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens

- 设计低帧率高维表示空间,提升符号可预测性
- 8Hz、768维下保持高重建质量与长序列稳定性
- 无需预训练模型,适合端到端语音生成场景
自回归语音生成中,序列长度、表征能力与长时稳定性难以兼顾。高帧率或高容量表示虽能保留更多细节,却易引发分布漂移与误差累积;而短且压缩的表示虽简化建模,又因带宽受限损失关键信息,限制重建精度与生成质量。本文提出一种低帧率、高维、高带宽的连续表示方案,通过协同设计编码器与流式生成框架,实现鲁棒的高保真重建、强单符号可预测性与优异的长时稳定性。为此,提出 Locodec 编码器,优化高维表示空间的可插值性与原生坐标的可辨识性,提升高维高带宽符号的预测能力;并提出 MP-ELD 框架,采用多路径信息路由与残差无分类器引导,缓解误差累积。在 8-Hz、768 维令牌下的实验表明,该方法保持高重建质量,提升单令牌可预测性,达成具有竞争力的 WER(词错误率),并在无外部 SSL/ASR 模型、预训练文本语言模型或后训练阶段的情况下,实现稳定长语音合成。
原文摘要 · Abstract (English)
Balancing sequence length, representational capacity, and long-horizon stability is a central problem in autoregressive (AR) speech and audio generation. Representations with higher frame rates or greater capacity can preserve more signal detail, but they also make streaming generation more vulnerable to distribution drift and AR error accumulation. Conversely, shorter and more compressed representations simplify AR modeling, but their limited bandwidth may discard important components and constrain the upper bound of reconstruction fidelity and generation quality. We ask whether a low-frame-rate, high-dimensional, high-bandwidth continuous representation can be co-designed with a streaming generation framework to support robust high-fidelity reconstruction, strong single-token predictability, and superior long-horizon stability. We decompose this goal into two coupled problems: what geometric and statistical properties a high-dimensional representation space should have, and how an AR continuous-token generator should be structured to resist error accumulation. Accordingly, we propose Locodec, a locally encoded codec that shapes its representation space to improve the interpolatability of a lower-dimensional core manifold and the identifiability of the native high-dimensional coordinates, thereby improving the predictability of high-dimensional high-bandwidth tokens. We also propose MP-ELD, a single-token AR flow-matching framework that uses multi-path information routing and residual classifier-free guidance to mitigate error accumulation. Experiments with 8-Hz, 768-dimensional tokens show that our design preserves reconstruction quality, improves single-token predictability, achieves competitive WER, and maintains stable long-form synthesis, without using external SSL/ASR models, pretrained text language models, or post-training stages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。