arXiv:2506.22023cs.SDcs.CL2025-06被引 5

动态分块预测让语音合成更快更准,适合实时应用。

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy

  • 用多标记预测训练动态分块机制,自适应调整生成跨度。
  • 测试集上语音清晰度提升72.27%,推理速度加快2.61倍。
  • 适合追求高效率与鲁棒性的下一代语音合成系统。

近年来,自回归(AR)语言模型已成为语音合成的主流方法,具备表现力强和可扩展训练的优点。然而,依赖逐词预测的传统AR语音合成模型在处理长语音序列时面临显著挑战,难以构建稳定的帧间注意力,导致延迟增加且合成质量下降,限制了其实时应用可行性。为此,我们提出一种新型动态分块自回归合成框架DCAR,旨在提升AR语音生成的效率与清晰度鲁棒性。DCAR通过多标记预测训练引入分块到帧的注意力机制,利用轻量级在线策略训练模块实现可变语境下的动态分块预测。该方法显著降低序列长度依赖性,同时保持高质量合成效果。全面实证评估表明,相较于传统逐词预测模型,DCAR在测试集上实现了高达72.27%的清晰度提升,并获得2.61倍的推理加速。此外,我们进行了深入分析,验证其作为下一代语音合成系统通用基础的潜力。

原文摘要 · Abstract (English)

Recently, autoregressive (AR) language models have emerged as a dominant approach in speech synthesis, offering expressive generation and scalable training. However, conventional AR speech synthesis models relying on the next-token prediction paradigm often encounter significant challenges when handling long speech sequences. These models often struggle to construct stable frame-to-frame attention, leading to increased latency and degraded synthesis quality, thereby limiting their feasibility for real-time applications. To address these limitations, we introduce a novel dynamic chunk-wise autoregressive synthesis framework, termed DCAR, designed to enhance both efficiency and intelligibility robustness in AR speech generation. DCAR introduces a chunk-to-frame attention mechanism through training with multi-token prediction, enabling dynamic chunk prediction in variable speech contexts using a lightweight module trained on-policy. DCAR dynamically adjusts the token prediction span, significantly reducing the sequence length dependency while obtaining high synthesis quality. Comprehensive empirical evaluations demonstrate that DCAR substantially outperforms traditional next-token prediction models, achieving up to 72.27% intelligibility improvement and 2.61x inference speedup simultaneously on the test set. Furthermore, we conduct comprehensive analysis to support it as a versatile foundation for next-generation speech synthesis systems.

语音合成自回归高效生成动态预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。