arXiv:2502.02942eess.AScs.SD2025-02ICLR被引 47

用语言模型提升语音增强,让合成语音更自然、更连贯。

GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling

  • 将语音增强转为条件语言建模任务,分步生成语义与声学特征。
  • 在多个基准数据集上显著提升语音质量与泛化能力。
  • 适合关注语音可懂度与音色一致性的研究者或开发者。

语义信息指通过词汇、短语及上下文关系传递的意义。人类能借助熟悉的语言模式和上下文线索,在嘈杂环境中重建不完整或被遮蔽的语音信号。然而,现有语音增强(SE)方法常忽略语音中蕴含的丰富语义信息,而这些信息对提升语音可懂度、说话人一致性及整体质量至关重要。为此,我们引入语言模型作为高效的语义学习器,提出一种面向语言模型的语音增强框架——GenSE。具体地,将语音增强视为条件语言建模任务,而非传统连续信号回归问题。通过预训练自监督模型将语音切分为语义标记,再利用定制单量化神经编解码器生成声学标记。为提高语言模型预测稳定性,提出分层建模方法,将干净语义标记与干净声学标记的生成分为两个阶段。此外,在声学标记生成阶段引入标记链提示机制,确保语音音色一致性。在基准数据集上的实验结果表明,该方法在语音质量与泛化能力方面均优于现有先进系统。

原文摘要 · Abstract (English)

Semantic information refers to the meaning conveyed through words, phrases, and contextual relationships within a given linguistic structure. Humans can leverage semantic information, such as familiar linguistic patterns and contextual cues, to reconstruct incomplete or masked speech signals in noisy environments. However, existing speech enhancement (SE) approaches often overlook the rich semantic information embedded in speech, which is crucial for improving intelligibility, speaker consistency, and overall quality of enhanced speech signals. To enrich the SE model with semantic information, we employ language models as an efficient semantic learner and propose a comprehensive framework tailored for language model-based speech enhancement, called \textit{GenSE}. Specifically, we approach SE as a conditional language modeling task rather than a continuous signal regression problem defined in existing works. This is achieved by tokenizing speech signals into semantic tokens using a pre-trained self-supervised model and into acoustic tokens using a custom-designed single-quantizer neural codec model. To improve the stability of language model predictions, we propose a hierarchical modeling method that decouples the generation of clean semantic tokens and clean acoustic tokens into two distinct stages. Moreover, we introduce a token chain prompting mechanism during the acoustic token generation stage to ensure timbre consistency throughout the speech enhancement process. Experimental results on benchmark datasets demonstrate that our proposed approach outperforms state-of-the-art SE systems in terms of speech quality and generalization capability.

语音增强语言模型语义建模音色一致

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。