arXiv:2507.12197cs.SDcs.AI2025-07被引 4

用多码本量化提升语音合成质量,保留韵律与音色细节

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations

  • 采用多码本残差量化编码,端到端训练增强语义分离
  • 双自回归结构与延迟多头机制,提升音质并加速推理
  • 适合追求高保真语音生成的研究者与开发者

文本转语音(TTS)在离散建模范式下取得新进展。现有自回归方法多依赖单码本表示,导致显著信息损失。即使使用流匹配等后处理技术,仍难以恢复细粒度特征(如语调变化、说话人特有音色),尤其在歌唱或音乐合成等复杂场景中表现不佳。本文提出QTTS框架,基于新型音频编解码器QDAC。QDAC的核心创新在于以ASR为基础的自回归网络与GAN联合端到端训练,实现可扩展的近无损压缩和优异的语义特征解耦。QTTS采用两种创新策略建模离散码:分层并行结构利用双自回归架构建模跨码本依赖关系,提升合成质量;延迟多头方法通过固定延迟的并行预测加速推理。实验表明,该框架在合成质量及表达内容保留方面优于基线。结果表明,通过多码本建模扩大压缩规模是实现高保真、通用语音与音频生成的可行方向。

原文摘要 · Abstract (English)

Text-to-speech (TTS) synthesis has seen renewed progress under the discrete modeling paradigm. Existing autoregressive approaches often rely on single-codebook representations, which suffer from significant information loss. Even with post-hoc refinement techniques such as flow matching, these methods fail to recover fine-grained details (e.g., prosodic nuances, speaker-specific timbres), especially in challenging scenarios like singing voice or music synthesis. We propose QTTS, a novel TTS framework built upon our new audio codec, QDAC. The core innovation of QDAC lies in its end-to-end training of an ASR-based auto-regressive network with a GAN, which achieves superior semantic feature disentanglement for scalable, near-lossless compression. QTTS models these discrete codes using two innovative strategies: the Hierarchical Parallel architecture, which uses a dual-AR structure to model inter-codebook dependencies for higher-quality synthesis, and the Delay Multihead approach, which employs parallelized prediction with a fixed delay to accelerate inference speed. Our experiments demonstrate that the proposed framework achieves higher synthesis quality and better preserves expressive content compared to baseline. This suggests that scaling up compression via multi-codebook modeling is a promising direction for high-fidelity, general-purpose speech and audio generation.

语音合成多码本自回归量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。