arXiv:2509.18806eess.AS2025-09

改进时频域声码器的幅值相位联合估计,提升音质稳定性。

Rethinking the joint estimation of magnitude and phase for time-frequency domain neural vocoders

  • 通过调整结构、引入先验知识、优化输出格式三策略改善联合估计
  • 使双流模型性能接近单流模型,解决严重性能下降问题
  • 适合研究语音合成与声码器架构优化的研究者

基于时频域的神经声码器在生成高保真音频方面表现优异,但其幅值与相位目标的联合预测机制尚不明确。本文以Vocos(单流)和APNet2(双流)为例进行分析,在大规模数据集上发现APNet2存在严重性能崩溃。为稳定其表现,提出三种简单有效的策略:在拓扑空间优化结构以增强信息交互;在源空间引入先验知识辅助生成;在输出空间改进损失函数与反向传播机制。实验表明,所提方法显著提升了APNet2的联合估计能力,缩小了单流与双流模型间的性能差距。

原文摘要 · Abstract (English)

Time-frequency (T-F) domain-based neural vocoders have shown promising results in synthesizing high-fidelity audio. Nevertheless, it remains unclear on the mechanism of effectively predicting magnitude and phase targets jointly. In this paper, we start from two representative T-F domain vocoders, namely Vocos and APNet2, which belong to the single-stream and dual-stream modes for magnitude and phase estimation, respectively. When evaluating their performance on a large-scale dataset, we accidentally observe severe performance collapse of APNet2. To stabilize its performance, in this paper, we introduce three simple yet effective strategies, each targeting the topological space, the source space, and the output space, respectively. Specifically, we modify the architectural topology for better information exchange in the topological space, introduce prior knowledge to facilitate the generation process in the source space, and optimize the backpropagation process for parameter updates with an improved output format in the output space. Experimental results demonstrate that our proposed method effectively facilitates the joint estimation of magnitude and phase in APNet2, thus bridging the performance disparities between the single-stream and dual-stream vocoders.

声码器时频域语音合成联合估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。