用音频补全实现歌声风格转换,保留原曲旋律更自然。
Serenade: A Singing Style Conversion Framework Based On Audio Infilling
- 用流匹配模型补全目标音谱,学习新风格特征。
- 循环训练分离源风格,提升转换保真度。
- 保留原始音高重合成,避免跑调,适合音乐创作。
我们提出Serenade,一种新型歌声风格转换框架。尽管歌手身份转换已有进展,但歌声风格转换仍属空白。主要挑战包括:建模目标风格、解耦源风格、保留源旋律。针对目标风格,采用音频补全任务,利用流匹配模型基于掩码音谱的补全信息与解耦声学特征进行预测。为解耦源风格,采用循环训练:以合成转换样本为输入,重建原始源音谱。为更好保留旋律,引入基于源-滤波器的后处理模块,使用原始基频(F0)模式重合成波形。实验表明,Serenade在泛化风格转换中表现最佳,尤其擅长处理气息型和混合型演唱风格。使用原始F0重合成有效缓解跑调问题并提升自然度,但因未改变基频模式,相似度略有下降。
原文摘要 · Abstract (English)
We propose Serenade, a novel framework for the singing style conversion (SSC) task. Although singer identity conversion has made great strides in the previous years, converting the singing style of a singer has been an unexplored research area. We find three main challenges in SSC: modeling the target style, disentangling source style, and retaining the source melody. To model the target singing style, we use an audio infilling task by predicting a masked segment of the target mel-spectrogram with a flow-matching model using the complement of the masked target mel-spectrogram along with disentangled acoustic features. On the other hand, to disentangle the source singing style, we use a cyclic training approach, where we use synthetic converted samples as source inputs and reconstruct the original source mel-spectrogram as a target. Finally, to retain the source melody better, we investigate a post-processing module using a source-filter-based vocoder and resynthesize the converted waveforms using the original F0 patterns. Our results showed that the Serenade framework can handle generalized SSC tasks with the best overall similarity score, especially in modeling breathy and mixed singing styles. We also found that resynthesizing with the original F0 patterns alleviated out-of-tune singing and improved naturalness, but found a slight tradeoff in similarity due to not changing the F0 patterns into the target style.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。