基于声源-滤波模型的高保真歌声合成,提升音高精度与自然度。
SiFiSinger: A High-Fidelity End-to-End Singing Voice Synthesizer based on Source-filter Model
- 用梅尔倒谱分离音高与频谱特征,减少预测误差
- 引入可微分的音高损失和倒谱损失,提升音高与语音包络精度
- 适合追求高音质歌声合成的研究者与音乐生成应用
本文提出一种基于声源-滤波机制的端到端歌声合成系统,直接将歌词与旋律线索转化为富有表现力且高保真的类人歌声。与VISinger 2类似,该系统借鉴VITS的训练范式,包含基频(F0)预测器和波形生成解码器。为解决梅尔谱特征与基频信息耦合导致的基频预测误差问题,本文提出两种策略:首先,采用梅尔倒谱(mcep)特征解耦梅尔谱与基频特性;其次,受神经声源-滤波模型启发,引入源激励信号作为基频在歌声合成中的表示,以更准确捕捉音高细节。同时,使用可微分的mcep损失和基频损失作为波形解码器的监督,强化语音包络与音高的预测准确性。在Opencpop数据集上的实验表明,所提模型在合成质量与音准精度方面均具显著优势。
原文摘要 · Abstract (English)
This paper presents an advanced end-to-end singing voice synthesis (SVS) system based on the source-filter mechanism that directly translates lyrical and melodic cues into expressive and high-fidelity human-like singing. Similarly to VISinger 2, the proposed system also utilizes training paradigms evolved from VITS and incorporates elements like the fundamental pitch (F0) predictor and waveform generation decoder. To address the issue that the coupling of mel-spectrogram features with F0 information may introduce errors during F0 prediction, we consider two strategies. Firstly, we leverage mel-cepstrum (mcep) features to decouple the intertwined mel-spectrogram and F0 characteristics. Secondly, inspired by the neural source-filter models, we introduce source excitation signals as the representation of F0 in the SVS system, aiming to capture pitch nuances more accurately. Meanwhile, differentiable mcep and F0 losses are employed as the waveform decoder supervision to fortify the prediction accuracy of speech envelope and pitch in the generated speech. Experiments on the Opencpop dataset demonstrate efficacy of the proposed model in synthesis quality and intonation accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。