arXiv:2602.06213eess.AS2026-02

用语言模型损失减少低比特率语音编码中的发音幻觉。

From Hallucination to Articulation: Language Model-Driven Losses for Ultra Low-Bitrate Neural Speech Coding

  • 引入语言模型驱动的损失函数,利用预训练模型对齐语音与文本。
  • 在极低比特率下,显著提升语义保真度,主观评价更接近原声。
  • 适合关注语音编码质量与语义保留的研究者或工程师。

低比特率的基于深度神经网络的语音编码中普遍存在『音素幻觉』(Phoneme Hallucinations, PH),即生成解码器在信息严重压缩的标记下尝试合成看似合理的输出。本文提出语言模型驱动的损失(LM loss),并证明其在超低比特率场景下优于语义蒸馏(SD)目标。所提方法基于预训练的语言模型,将语音与文本关联。当缺乏真实转录时,通过修改流行的自动语音识别模型Whisper,将解码语音与输入语音的ASR推断转录进行比较;否则,采用时序文本正则化(TTR),对比解码语音的WavLM表示与真实转录的BERT表示。在参考编码器上测试,该三阶段训练框架基于多个主流编码器设计。主观与客观评估表明,LM损失能更强地引导从自监督语音表示中提取语义信息,显著提升人类感知的语义一致性,同时保持整体输出质量。演示样本、代码和模型检查点已公开。

原文摘要 · Abstract (English)

``Phoneme Hallucinations (PH)'' commonly occur in low-bitrate DNN-based codecs. It is the generative decoder's attempt to synthesize plausible outputs from excessively compressed tokens missing some semantic information. In this work, we propose language model-driven losses (LM loss) and show they may alleviate PHs better than a semantic distillation (SD) objective in very-low-bitrate settings. The proposed LM losses build upon language models pretrained to associate speech with text. When ground-truth transcripts are unavailable, we propose to modify a popular automatic speech recognition (ASR) model, Whisper, to compare the decoded utterance against the ASR-inferred transcriptions of the input speech. Else, we propose to use the timed-text regularizer (TTR) to compare WavLM representations of the decoded utterance against BERT representations of the ground-truth transcriptions. We test and compare LM losses against an SD objective, using a reference codec whose three-stage training regimen was designed after several popular codecs. Subjective and objective evaluations conclude that LM losses may provide stronger guidance to extract semantic information from self-supervised speech representations, boosting human-perceived semantic adherence while preserving overall output quality. Demo samples, code, and checkpoints are available online.

语音编码语言模型低比特率语义保真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。