arXiv:2505.24496eess.AS2025-05被引 10

通过压缩长序列语音标记提升生成质量

Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation

  • 保留局部上下文并压缩远距离信息以减少冗余
  • 在多个编码器和模型上验证,显著提升生成效果
  • 适合追求高保真语音生成的研究者与工程师

神经音频编解码器作为语音标记生成器,在语音生成领域展现出巨大潜力。然而,为保证高保真音频重建,神经音频编解码器通常将语音编码为较长的语音标记序列,给下游语言模型的长上下文建模带来挑战。我们观察到语音标记序列具有短程依赖性:在文本到语音任务中,语音与文本存在单调对齐,当前标记的预测主要依赖其局部上下文,而远距离标记对当前预测贡献较小且常含冗余信息。受此启发,我们提出一种‘从压缩到精细’的语言建模方法,解决神经编解码器语言模型中长序列语音标记的难题:(1) 精细初始与短程信息:预测时保留提示和局部标记,确保文本对齐与副语言信息完整性;(2) 压缩远距离上下文:将远距离标记段压缩为紧凑表示,降低冗余信息同时保留关键语义。在多种神经音频编解码器和下游语言模型上的大量实验验证了该方法的有效性与通用性,突显了在神经编解码器语言模型中进行标记压缩的重要性。音频样例演示将发布于 https://anonymous.4open.science/r/SpeechTokenPredictionViaCompressedToFinedLM。

原文摘要 · Abstract (English)

Neural audio codecs, used as speech tokenizers, have demonstrated remarkable potential in the field of speech generation. However, to ensure high-fidelity audio reconstruction, neural audio codecs typically encode audio into long sequences of speech tokens, posing a significant challenge for downstream language models in long-context modeling. We observe that speech token sequences exhibit short-range dependency: due to the monotonic alignment between text and speech in text-to-speech (TTS) tasks, the prediction of the current token primarily relies on its local context, while long-range tokens contribute less to the current token prediction and often contain redundant information. Inspired by this observation, we propose a \textbf{compressed-to-fine language modeling} approach to address the challenge of long sequence speech tokens within neural codec language models: (1) \textbf{Fine-grained Initial and Short-range Information}: Our approach retains the prompt and local tokens during prediction to ensure text alignment and the integrity of paralinguistic information; (2) \textbf{Compressed Long-range Context}: Our approach compresses long-range token spans into compact representations to reduce redundant information while preserving essential semantics. Extensive experiments on various neural audio codecs and downstream language models validate the effectiveness and generalizability of the proposed approach, highlighting the importance of token compression in improving speech generation within neural codec language models. The demo of audio samples will be available at https://anonymous.4open.science/r/SpeechTokenPredictionViaCompressedToFinedLM.

语音生成语言模型编码压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。