用语言模型提升语音分离后的可懂度和连贯性
SLM-SS: Speech Language Model for Generative Speech Separation
- 将语音分离建模为离散多码本序列生成任务
- 在LibriMix上显著提升语音可懂度,下游任务表现更优
- 兼顾自回归与非自回归解码,平衡效果与效率
语音分离(SS)在神经网络方法推动下,信号级指标显著提升,但分离后语音的可懂度常受影响,进而损害下游语音识别等任务性能。本文提出SLM-SS,首次将语音语言模型引入语音分离,旨在增强分离信号的可懂度与连贯性。我们将语音分离建模为离散多码本序列生成问题,采用编码器-解码器结构将量化语音混合信号映射为目标词元。除自回归建模外,还引入非自回归模型以提升残差词元的解码效率。在LibriMix数据集上的实验表明,该方法显著改善了语音可懂度,提升了多种下游任务的语言一致性,优于现有方法。
原文摘要 · Abstract (English)
Speech separation (SS) has advanced significantly with neural network-based methods, showing improved performance on signal-level metrics. However, these methods often struggle to maintain speech intelligibility in the separated signals, which can negatively affect the performance of downstream tasks such as speech recognition. In this work, we propose SLM-SS, a novel approach that applies speech language models to SS, aiming to enhance the intelligibility and coherence of the separated signals. We frame SS as discrete multi-codebook sequence generation, using Encoder-Decoder models to map quantized speech mixtures to target tokens. In addition to the autoregressive modeling strategy, we introduce a non-autoregressive model to improve decoding efficiency for residual tokens. Experimental results on the LibriMix dataset demonstrate that our approach shows significantly better preservation of speech intelligibility, leading to improved linguistic consistency in a variety of downstream tasks compared to existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。