arXiv:2609.03940cs.SDcs.AI2026-09

用连续音频编码表示改进语音增强,兼顾质量与效率

Masked Autoregressive Speech Enhancement with Continuous Neural Audio Codec Representations

论文配图:Masked Autoregressive Speech Enhancement with Continuous Neural Audio Codec Representations
图 1 · 摘自论文原文
  • 基于连续编码的自回归迭代解码,提升语音重建精度
  • 在相同模型和训练条件下,实现性能与计算成本灵活权衡
  • 适合关注语音质量与推理效率平衡的研究者

以往基于掩码生成建模的语音增强方法多依赖神经音频编码器(NAC)产生的离散音素表示。最近研究表明,使用NAC的连续潜在表示在语音质量和可懂度上更具优势。本文提出掩码自回归语音增强(MARSE),通过迭代解码掩码后的纯净语音帧,利用连续NAC表示进行语音增强。我们考察了多种解码策略,在保持相同深度神经网络(Conformer模型)、相同编码器(DAC编码器)及相同训练设置的前提下进行比较。结果表明,MARSE可在语音增强性能与计算开销之间实现灵活权衡。音频示例与代码已公开。

原文摘要 · Abstract (English)

Most previous work on speech enhancement (SE) based on masked generative modeling relied on discrete token representations of audio signals, obtained using neural audio codecs (NACs). However, a recent study has shown that continuous latent representations of NACs can be advantageous for SE in terms of speech quality and intelligibility. In this work, we propose masked autoregressive SE (MARSE), a method for SE based on iterative decoding of masked clean speech frames using continuous NAC representations of speech. In particular, we investigate a set of different decoding policies, ceteris paribus, that is, using the same DNN (a Conformer model), the same NAC (the DAC codec) and the same training setup. The results show that MARSE enables a flexible trade-off between SE performance and computational cost. Audio examples and code are available online.

语音增强连续编码自回归模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。