arXiv:2502.16240eess.AScs.AI2025-02中稿 · ICASSP 2025被引 20

用预训练音频编码器的连续嵌入做语音增强,效率高且效果好。

Speech Enhancement Using Continuous Embeddings of Neural Audio Codec

  • 在预训练NAC的连续嵌入空间中进行语音增强,避免离散标记处理
  • 实时因子仅0.005,GMAC为3.94,复杂度降低18倍
  • 适合云环境下的音频压缩传输场景,轻量高效

近年来,神经音频编码器(NAC)模型推动了其在语音处理任务中的应用,包括语音增强(SE)。本文提出一种新颖高效的SE方法,利用预训练NAC编码器的预量化输出——连续嵌入空间。与以往基于NAC的SE方法不同,这些方法依赖语言模型处理离散语音标记,而本工作在高度压缩的时间维度连续嵌入空间中直接进行语音增强。所提出的轻量级SE模型通过嵌入级损失优化,在性能上可媲美在更大数据集上训练的基线方法,同时实现0.005的极低实时因子。此外,在模拟云端音频传输环境中,其GMAC仅为3.94,相比Sepformer降低18倍。该研究展示了一种高效、适用于云端音频压缩传输的新NAC-based SE方案。

原文摘要 · Abstract (English)

Recent advancements in Neural Audio Codec (NAC) models have inspired their use in various speech processing tasks, including speech enhancement (SE). In this work, we propose a novel, efficient SE approach by leveraging the pre-quantization output of a pretrained NAC encoder. Unlike prior NAC-based SE methods, which process discrete speech tokens using Language Models (LMs), we perform SE within the continuous embedding space of the pretrained NAC, which is highly compressed along the time dimension for efficient representation. Our lightweight SE model, optimized through an embedding-level loss, delivers results comparable to SE baselines trained on larger datasets, with a significantly lower real-time factor of 0.005. Additionally, our method achieves a low GMAC of 3.94, reducing complexity 18-fold compared to Sepformer in a simulated cloud-based audio transmission environment. This work highlights a new, efficient NAC-based SE solution, particularly suitable for cloud applications where NAC is used to compress audio before transmission. Copyright 20XX IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works.

语音增强NAC嵌入空间云端应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。