用离散声学标记去噪提升大模型零样本语音合成的抗噪能力
Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising
- 通过神经编解码器设计,从噪声音频中恢复干净的声学标记
- 在噪声环境下合成语音质量显著优于现有方法,峰值信噪比提升1.8dB
- 适合需要高鲁棒性语音合成的场景,如语音助手、无障碍通信
基于大语言模型(LLM)的零样本文本到语音(TTS)方法常会保留音频提示中的声学环境,导致当提示音频含噪声时,合成语音质量下降。本文提出一种新型神经编解码器语音去噪器,并与先进LLM-TTS模型LauraTTS集成,实现抗噪零样本语音合成。该编解码器去噪器由音频编解码器、标记去噪器和嵌入重构器组成。标记去噪器从含噪标记中预测前两组干净声学标记,可作为LauraTTS的声学提示以生成高质量个性化语音,或经嵌入重构器和编解码器解码器转换为纯净语音波形。实验表明,所提编解码器去噪器性能优于当前最优语音增强(SE)方法,且集成后的噪声鲁棒LauraTTS在无需额外语音增强模型的情况下超越对比方法。
原文摘要 · Abstract (English)
Large language model (LLM) based zero-shot text-to-speech (TTS) methods tend to preserve the acoustic environment of the audio prompt, leading to degradation in synthesized speech quality when the audio prompt contains noise. In this paper, we propose a novel neural codec-based speech denoiser and integrate it with the advanced LLM-based TTS model, LauraTTS, to achieve noise-robust zero-shot TTS. The proposed codec denoiser consists of an audio codec, a token denoiser, and an embedding refiner. The token denoiser predicts the first two groups of clean acoustic tokens from the noisy ones, which can serve as the acoustic prompt for LauraTTS to synthesize high-quality personalized speech or be converted to clean speech waveforms through the embedding refiner and codec decoder. Experimental results show that our proposed codec denoiser outperforms state-of-the-art speech enhancement (SE) methods, and the proposed noise-robust LauraTTS surpasses the approach using additional SE models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。