用小模型+高效求解,快速高保真还原语音频谱相位。
Efficient Neural and Numerical Methods for High-Quality Online Speech Spectrogram Inversion via Gradient Theorem
- 8k参数小模型直接预测相位导数,大幅降低计算量。
- 增加1个时隙延迟,使神经网络推理成本减半。
- 利用三对角矩阵特性,线性复杂度求解,速度提升数个数量级。
近期在线语音频谱逆问题研究结合深度学习与梯度定理,直接从幅度预测相位导数,并通过最小二乘法估计相位,实现高质量重建。本文提出三项创新:首先,设计仅含8k参数的新型神经网络架构,仅为先前最优方法的1/30;其次,增加1个时隙延迟,使神经网络推理成本进一步减半;第三,观察到最小二乘问题中系数矩阵为三对角阵,提出一种利用其三对角结构与半正定性的线性复杂度求解器,实现数个数量级的速度提升。相关音频样本已公开。
原文摘要 · Abstract (English)
Recent work in online speech spectrogram inversion effectively combines Deep Learning with the Gradient Theorem to predict phase derivatives directly from magnitudes. Then, phases are estimated from their derivatives via least squares, resulting in a high quality reconstruction. In this work, we introduce three innovations that drastically reduce computational cost, while maintaining high quality: Firstly, we introduce a novel neural network architecture with just 8k parameters, 30 times smaller than previous state of the art. Secondly, increasing latency by 1 hop size allows us to further halve the cost of the neural inference step. Thirdly, we we observe that the least squares problem features a tridiagonal matrix and propose a linear-complexity solver for the least squares step that leverages tridiagonality and positive-semidefiniteness, achieving a speedup of several orders of magnitude. We release samples online.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。