通过高频音高分解实现歌声颤音精准控制,提升转换自然度。
VibE-SVC: Vibrato Extraction with High-frequency F0 Contour for Singing Voice Conversion
- 用离散小波变换分离音高轮廓的频率成分,显式提取颤音特征。
- 主观与客观评估均证明转换后歌声保真度高、风格可控性强。
- 适合需要精细调节歌唱情感表达的研究者或音乐创作应用。
控制歌唱风格对实现富有表现力和自然的歌声至关重要。在多种风格因素中,颤音在传达情感和增强音乐深度方面起着关键作用。然而,由于其动态特性,建模颤音仍具挑战性,导致在歌声转换中难以控制。为此,我们提出VibE-SVC,一种可调控的歌声转换模型,通过离散小波变换显式提取并操控颤音。与以往隐式建模颤音的方法不同,本方法将音高轮廓分解为频率成分,实现精确转移,从而增强风格控制灵活性。实验结果表明,VibE-SVC能有效转换歌唱风格,同时保持说话人相似性。主观与客观评估均证实其转换质量优异。
原文摘要 · Abstract (English)
Controlling singing style is crucial for achieving an expressive and natural singing voice. Among the various style factors, vibrato plays a key role in conveying emotions and enhancing musical depth. However, modeling vibrato remains challenging due to its dynamic nature, making it difficult to control in singing voice conversion. To address this, we propose VibESVC, a controllable singing voice conversion model that explicitly extracts and manipulates vibrato using discrete wavelet transform. Unlike previous methods that model vibrato implicitly, our approach decomposes the F0 contour into frequency components, enabling precise transfer. This allows vibrato control for enhanced flexibility. Experimental results show that VibE-SVC effectively transforms singing styles while preserving speaker similarity. Both subjective and objective evaluations confirm high-quality conversion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。