用优化算法提升从梅尔频谱重建语音和音效的精度与速度。
Mel-Spectrogram Inversion via Alternating Direction Method of Multipliers
- 基于ADMM的联合估计方法,同步优化幅度与相位。
- 实验表明在语音和音效重建上优于现有方法。
- 适合需要高质量音频重建的研究与开发人员。
从梅尔频谱重建时域信号称为梅尔频谱反演,广泛应用于语音和环境音效合成。本文提出一种基于交替方向乘子法(ADMM)的梅尔频谱反演方法,旨在通过联合估计全频段短时傅里叶变换(STFT)的幅度与相位来提升重建质量。传统级联式方法存在误差累积问题,而联合估计虽更优但仍需大量迭代。本文利用ADMM在非凸优化中的成功经验,设计出高效变量更新策略,充分利用变量间的条件独立性。实验结果表明,该方法在语音和环境音效重建任务中均表现出色,显著提升了重建效果。
原文摘要 · Abstract (English)
Signal reconstruction from its mel-spectrogram is known as mel-spectrogram inversion and has many applications, including speech and foley sound synthesis. In this paper, we propose a mel-spectrogram inversion method based on a rigorous optimization algorithm. To reconstruct a time-domain signal with inverse short-time Fourier transform (STFT), both full-band STFT magnitude and phase should be predicted from a given mel-spectrogram. Their joint estimation has outperformed the cascaded full-band magnitude prediction and phase reconstruction by preventing error accumulation. However, the existing joint estimation method requires many iterations, and there remains room for performance improvement. We present an alternating direction method of multipliers (ADMM)-based joint estimation method motivated by its success in various nonconvex optimization problems including phase reconstruction. An efficient update of each variable is derived by exploiting the conditional independence among the variables. Our experiments demonstrate the effectiveness of the proposed method on speech and foley sounds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。