arXiv:2601.06006eess.AScs.SD2026-01

用生成模型提升语音提取质量,兼顾清晰度与保真度。

Discriminative-Generative Target Speaker Extraction with Decoder-Only Language Models

  • 先用判别模型抑制干扰,再用生成模型重建高质量语音
  • 在多个基准上同时提升语音自然度、可懂度和说话人一致性
  • 适合需要高保真语音输出的应用场景

目标说话人提取(TSE)旨在从混合语音中恢复指定说话人的语音,而语音增强(SE)则关注在噪声环境下提升语音质量。现有TSE与SE系统多基于判别式建模,虽有较强的干扰抑制能力,但感知质量与自然度仍受限。为此,本文提出LauraTSE,一个基于自回归解码器的生成式TSE模型。尽管生成式建模在质量提升上有潜力,但纯生成式方法在复杂声学条件下易出现幻觉、内容漂移和控制性差的问题。因此,我们设计了判别-生成两阶段框架:判别前端首先生成强干扰抑制的目标相关表示,生成后端在神经音频编解码空间中重建高质量语音。该设计融合了判别模型的可控性与生成模型的重构能力。进一步研究了多种协作策略,包括前端冻结、联合微调、SI-SDR正则化以及自回归/非自回归推理。在TSE与SE基准上的实验表明,所提框架在感知质量、可懂度与说话人一致性之间取得更好平衡,优于纯判别或纯生成基线。

原文摘要 · Abstract (English)

Target speaker extraction (TSE) aims to recover the speech of a desired speaker from a mixture given a short enrollment utterance, while speech enhancement (SE) focuses on improving speech quality under noisy conditions. Most existing TSE and SE systems are based on discriminative modeling and have shown strong interference suppression ability, but they often remain limited in perceptual quality and naturalness. To address this issue, we first introduce LauraTSE, a generative TSE model built on an autoregressive decoder-only language model. Although generative modeling is promising for quality enhancement, purely generative TSE may suffer from hallucination, content drift, and limited controllability in complex acoustic conditions. We therefore propose a discriminative-generative two-stage framework, where a discriminative front-end first produces target-related representations with strong interference suppression, and a generative back-end then reconstructs high-quality speech in the neural audio codec representation space. This design combines the controllability of discriminative extraction with the reconstruction capability of generative modeling. We further investigate several collaboration strategies for the two-stage framework, including front-end freezing, joint fine-tuning, SI-SDR regularization, and autoregressive/non-autoregressive inference. Experimental results on both TSE and SE benchmarks show that the proposed framework achieves a better balance among perceptual quality, intelligibility, and speaker consistency than purely discriminative or purely generative baselines.

语音提取生成模型语音增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。