arXiv:2511.07116cs.SD2025-11被引 2

将语音合成视为音频修复任务,用扩散模型实现高效高质语音生成。

BridgeVoC: Revitalizing Neural Vocoder from a Restoration Perspective

  • 把声谱重建看作特定音频修复,用薛定谔桥框架连接目标与退化声谱。
  • 仅需4步采样即达顶尖性能,单步推理仍优于现有生成模型。
  • 设计分频带卷积注意力网络,支持高效时频上下文建模,参数更少。

本文从音频修复视角重审神经声码器任务,提出新型扩散声码器BridgeVoC。通过秩分析发现梅尔谱与其他常见声学退化因素具有相似秩特性,将声码器任务建模为特定音频修复问题,其中目标谱的范围空间谱(RSS)作为退化输入。基于此,引入薛定谔桥框架,将RSS与目标谱定义为随机生成轨迹的两端。为进一步利用时频域子带的层次先验,设计新颖的分频带感知卷积扩散网络:子带按非均匀策略划分,采用大核卷积式注意力模块实现高效时频上下文建模。为支持单步推理,提出全向蒸馏损失,结合目标相关性与双射一致性损失,促进教师到学生模型的有效信息传递。在多个基准和分布外数据集上进行综合实验。定量与定性结果表明,相比现有基于GAN、DDPM及流匹配的先进基线,BridgeVoC以更少参数、更低计算开销和相当推理速度,仅需4步采样即达到顶尖性能;单步推理下依然保持一致优势。

原文摘要 · Abstract (English)

This paper revisits the neural vocoder task through the lens of audio restoration and propose a novel diffusion vocoder called BridgeVoC. Specifically, by rank analysis, we compare the rank characteristics of Mel-spectrum with other common acoustic degradation factors, and cast the vocoder task as a specialized case of audio restoration, where the range-space spectral (RSS) surrogate of the target spectrum acts as the degraded input. Based on that, we introduce the Schrodinger bridge framework for diffusion modeling, which defines the RSS and target spectrum as dual endpoints of the stochastic generation trajectory. Further, to fully utilize the hierarchical prior of subbands in the time-frequency (T-F) domain, we elaborately devise a novel subband-aware convolutional diffusion network as the data predictor, where subbands are divided following an uneven strategy, and convolutional-style attention module is employed with large kernels for efficient T-F contextual modeling. To enable single-step inference, we propose an omnidirectional distillation loss to facilitate effective information transfer from the teacher model to the student model, and the performance is improved by combining target-related and bijective consistency losses. Comprehensive experiments are conducted on various benchmarks and out-of-distribution datasets. Quantitative and qualitative results show that while enjoying fewer parameters, lower computational cost, and competitive inference speed, the proposed BridgeVoC yields stateof-the-art performance over existing advanced GAN-, DDPMand flow-matching-based baselines with only 4 sampling steps. And consistent superiority is still achieved with single-step inference.

语音生成扩散模型声码器音频修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。