arXiv:2503.00493eess.AScs.AI2025-03ACL被引 39

用大模型提升语音增强的通用性,兼顾音质与跨任务适应能力

LLaSE-G1: Incentivizing Generalization Capability for LLaMA-based Speech Enhancement

  • 用WavLM连续特征输入,结合X-Codec2语音编码,保留原始声学信息
  • 双通道设计统一多种语音增强任务,无需任务标识符即可泛化
  • 在未见过的任务上展现新能力,测试时扩展效果优于传统方法

近期语言模型在语义理解与上下文建模方面取得显著进展,已成功应用于生成式语音增强。然而,多数基于语言模型的语音增强方法过于关注语义信息,忽视了声学信息的关键作用,导致增强后音频存在声学不一致问题,且跨任务泛化能力有限。本文提出LLaSE-G1,一种基于LLaMA的语言模型,旨在增强语音增强的泛化能力。首先,为缓解声学不一致,LLaSE-G1采用WavLM的连续表示作为输入,预测X-Codec2的语音标记,最大化声学保真度。其次,为提升泛化能力,引入双通道输入输出结构,统一多个语音增强任务,无需任务特定标识符。第三,实验表明,LLaSE-G1优于先前的任务专用判别式与生成式语音增强模型,在测试时展现缩放效应,并具备处理未见任务的涌现能力。此外,我们开源代码与模型,以支持该领域的进一步研究。

原文摘要 · Abstract (English)

Recent advancements in language models (LMs) have demonstrated strong capabilities in semantic understanding and contextual modeling, which have flourished in generative speech enhancement (SE). However, many LM-based SE approaches primarily focus on semantic information, often neglecting the critical role of acoustic information, which leads to acoustic inconsistency after enhancement and limited generalization across diverse SE tasks. In this paper, we introduce LLaSE-G1, a LLaMA-based language model that incentivizes generalization capabilities for speech enhancement. LLaSE-G1 offers the following key contributions: First, to mitigate acoustic inconsistency, LLaSE-G1 employs continuous representations from WavLM as input and predicts speech tokens from X-Codec2, maximizing acoustic preservation. Second, to promote generalization capability, LLaSE-G1 introduces dual-channel inputs and outputs, unifying multiple SE tasks without requiring task-specific IDs. Third, LLaSE-G1 outperforms prior task-specific discriminative and generative SE models, demonstrating scaling effects at test time and emerging capabilities for unseen SE tasks. Additionally, we release our code and models to support further research in this area.

语音增强大模型泛化能力声学保真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。