提出可预测的低比特语音编码器,兼顾高保真重建与自回归模型生成效率。
ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure

- 通过保留预量化语音的音素结构,提升自回归预测性
- 在650和800 bps下实现更高重建质量与可预测性平衡
- 适用于需要清晰语音表示的文本到语音合成任务
神经语音编解码器在语言模型时代面临根本矛盾:支持高保真重建的词元未必易于自回归模型预测。对多种编解码器与自监督语音表征的受控分析表明,量化前更清晰的音素结构始终与更易预测的词元相关。然而,仅靠音素结构不足以实现高保真重建,还需重建相关的声学细节。基于此,我们提出ReLMCodec,一种基于‘保持-控制-精炼’原则的低比特单码本语音编解码器:它在量化器输入处保留冻结自监督学习(SSL)特征的语言组织结构,通过预量化锚点保持适应(PAPA)控制重建驱动的偏移,再利用仅训练阶段的WavLM-Large L24教师模型精炼量化潜空间,减少音素级词元碎片化。上述组件使声学细节支持波形重建,同时保持词元序列对自回归模型的可预测性。在650和800 bps下,我们的评估显示ReLMCodec实现了更高的单流预测性-重建性能前沿,且该优势在下游文本到语音(TTS)合成中体现于可懂度与说话人相似性提升。
原文摘要 · Abstract (English)
Neural speech codecs face a fundamental tension in the language-model era: tokens that support high-fidelity reconstruction are not necessarily easy for autoregressive models to predict. Our controlled analysis of diverse codec and self-supervised speech representations shows that clearer phoneme structure before discrete code assignment is consistently associated with easier autoregressive token prediction. Yet phoneme structure alone is insufficient for high-fidelity reconstruction, which also requires reconstruction-relevant acoustic detail. Guided by this observation, we introduce ReLMCodec, a low-bitrate single-codebook speech codec built upon a preserve--control--refine principle: it preserves the linguistic organization of frozen self-supervised learning (SSL) features at the quantizer input, controls reconstruction-driven drift through Pre-quantization Anchor-Preserving Adaptation (PAPA), and refines the quantized latent space with a training-only WavLM-Large L24 teacher to reduce phoneme-level token fragmentation. Together, these components allow acoustic detail to support waveform reconstruction while keeping the resulting token sequence predictable for autoregressive models. At 650 and 800 bps, ReLMCodec moves the empirical single-stream predictability--reconstruction frontier in our evaluations, with gains that carry over to downstream text-to-speech (TTS) synthesis in both intelligibility and speaker similarity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。