用神经音频编码的离散特征恢复语音高频细节,提升清晰度和自然度。
Vector Quantized Diffusion Model Based Speech Bandwidth Extension
- 基于向量量化扩散模型,从压缩离散码本中重建高频语音
- 在谱距离和ViSQOL指标上均显著优于现有方法
- 适合语音增强、低码率通信场景应用
神经音频编码(NAC)的进展为音频信号处理带来了新可能。越来越多研究尝试利用NAC的潜在特征完成各类语音处理任务。本文首次提出一种基于NAC离散特征的语音带宽扩展(BWE)方法,通过在高度压缩的离散令牌中恢复高频细节,显著提升了语音的可懂度与自然度。所提框架基于向量量化扩散模型,融合先进NAC、扩散模型与Mamba-2,有效重建高频语音成分。大量实验表明,该方法在对数谱距和ViSQOL指标上均表现优异,显著改善语音质量。
原文摘要 · Abstract (English)
Recent advancements in neural audio codec (NAC) unlock new potential in audio signal processing. Studies have increasingly explored leveraging the latent features of NAC for various speech signal processing tasks. This paper introduces the first approach to speech bandwidth extension (BWE) that utilizes the discrete features obtained from NAC. By restoring high-frequency details within highly compressed discrete tokens, this approach enhances speech intelligibility and naturalness. Based on Vector Quantized Diffusion, the proposed framework combines the strengths of advanced NAC, diffusion models, and Mamba-2 to reconstruct high-frequency speech components. Extensive experiments demonstrate that this method exhibits superior performance across both log-spectral distance and ViSQOL, significantly improving speech quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。