arXiv:2603.07285eess.AScs.LG2026-03被引 1

用轻量级网络实现8-48kHz音频带宽扩展,高保真且超快

Fast and Flexible Audio Bandwidth Extension via Vocos

  • 基于Vocos的神经声码器骨干+轻量级交叉滤波器,支持任意上采样倍数
  • 在A100上实时因子达0.0001,在8核CPU上为0.0053,性能极优
  • 适合低延迟语音通信与嵌入式音频增强场景

我们提出一种基于Vocos的带宽扩展模型,通过生成缺失的高频内容,将8-48 kHz音频信号进行增强。输入信号重采样至48 kHz后由神经声码器主干处理,使单一网络可支持任意上采样比例。一个受Linkwitz-Riley启发的轻量级修正模块,通过平滑交叉方式融合原始低频与生成的高频成分。验证结果显示,该模型在对数谱距离上表现具有竞争力,同时在NVIDIA A100 GPU上实现实时因子0.0001,在8核CPU上为0.0053,展现出极高吞吐下的高质量带宽扩展能力。

原文摘要 · Abstract (English)

We propose a Vocos-based bandwidth extension model that enhances audio at 8-48 kHz by generating missing high-frequency content. Inputs are resampled to 48 kHz and processed by a neural vocoder backbone, enabling a single network to support arbitrary upsampling ratios. A lightweight Linkwitz-Riley-inspired refiner merges the original low band with the generated high frequencies via a smooth crossover. On validation, the model achieves competitive log-spectral distance while running at a real-time factor of 0.0001 on an NVIDIA A100 GPU and 0.0053 on an 8-core CPU, demonstrating practical, high-quality BWE at extreme throughput.

音频增强神经声码器带宽扩展实时处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。