arXiv:2507.12825cs.SDcs.LG2025-07被引 3

用声学离散标记实现自回归语音增强,提升保真度与可懂度。

Autoregressive Speech Enhancement via Acoustic Tokens

  • 采用声学令牌与自回归架构,建模语音时序依赖关系。
  • 在VoiceBank和Libri1Mix上优于语义令牌,保留说话人特征。
  • 适合关注语音保真与多模态融合的研究者。

在语音处理流程中,提升真实录音的质量与可懂度至关重要。尽管监督回归是语音增强的主流方法,但音频分词正成为与多模态融合的有前景替代方案。然而,基于离散表示的语音增强研究仍有限。先前工作多聚焦于语义令牌,往往丢失关键声学细节(如说话人身份)。此外,这些研究通常采用非自回归模型,假设输出条件独立,忽视了自回归建模的潜力。为弥补这些空白,本文:1)全面研究声学令牌在语音增强中的表现,涵盖码率与噪声强度的影响;2)提出一种专为该任务设计的基于转换器的自回归架构。在VoiceBank与Libri1Mix数据集上的实验表明,声学令牌在保留说话人身份方面优于语义令牌,且自回归方法可进一步提升性能。然而,离散表示仍不及连续表示,凸显该领域尚需深入研究。

原文摘要 · Abstract (English)

In speech processing pipelines, improving the quality and intelligibility of real-world recordings is crucial. While supervised regression is the primary method for speech enhancement, audio tokenization is emerging as a promising alternative for a smooth integration with other modalities. However, research on speech enhancement using discrete representations is still limited. Previous work has mainly focused on semantic tokens, which tend to discard key acoustic details such as speaker identity. Additionally, these studies typically employ non-autoregressive models, assuming conditional independence of outputs and overlooking the potential improvements offered by autoregressive modeling. To address these gaps we: 1) conduct a comprehensive study of the performance of acoustic tokens for speech enhancement, including the effect of bitrate and noise strength; 2) introduce a novel transducer-based autoregressive architecture specifically designed for this task. Experiments on VoiceBank and Libri1Mix datasets show that acoustic tokens outperform semantic tokens in terms of preserving speaker identity, and that our autoregressive approach can further improve performance. Nevertheless, we observe that discrete representations still fall short compared to continuous ones, highlighting the need for further research in this area.

语音增强自回归声学令牌

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。