arXiv:2507.19062cs.SDeess.AS2025-07被引 4

用分层语言模型统一处理多种语音畸变,提升真实场景下的语音增强效果。

From Continuous to Discrete: Cross-Domain Collaborative General Speech Enhancement via Hierarchical Language Models

  • 分两阶段处理:先增强连续特征,再用语言模型生成离散语音标记
  • 在复合畸变场景下优于现有模型,尤其在噪声与混响共存时表现突出
  • 适合需要鲁棒语音增强的智能语音系统开发者使用

本文提出OmniGSE框架,旨在应对真实场景中语音信号面临的多种畸变,包括背景噪声、混响、带宽限制、信号削波和网络丢包。现有方法多针对单一畸变优化,难以有效处理多重畸变共存的情况。OmniGSE通过两阶段架构融合判别式与生成式方法,实现跨域协同优化:第一阶段采用轻量级通道分离的NAC-RoFormer增强连续特征;第二阶段利用分层语言模型生成离散标记以重建高质量语音。该模型包含根语言模型(RootLM)和多个分支语言模型(BranchLM),RootLM建模跨码本层的通用声学特征,BranchLM显式捕捉不同码本层级间的渐进关系。实验表明,OmniGSE在多个基准测试中超越现有模型,尤其在复合畸变场景下表现优异,展现出在真实应用中实现强健、通用语音增强的巨大潜力。

原文摘要 · Abstract (English)

This paper introduces OmniGSE, a novel general speech enhancement (GSE) framework designed to mitigate the diverse distortions that speech signals encounter in real-world scenarios. These distortions include background noise, reverberation, bandwidth limitations, signal clipping, and network packet loss. Existing methods typically focus on optimizing for a single type of distortion, often struggling to effectively handle the simultaneous presence of multiple distortions in complex scenarios. OmniGSE bridges this gap by integrating the strengths of discriminative and generative approaches through a two-stage architecture that enables cross-domain collaborative optimization. In the first stage, continuous features are enhanced using a lightweight channel-split NAC-RoFormer. In the second stage, discrete tokens are generated to reconstruct high-quality speech through language models. Specifically, we designed a hierarchical language model structure consisting of a RootLM and multiple BranchLMs. The RootLM models general acoustic features across codebook layers, while the BranchLMs explicitly capture the progressive relationships between different codebook levels. Experimental results demonstrate that OmniGSE surpasses existing models across multiple benchmarks, particularly excelling in scenarios involving compound distortions. These findings underscore the framework's potential for robust and versatile speech enhancement in real-world applications.

语音增强分层模型多畸变

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。