arXiv:2409.08039cs.SDeess.AS2024-09被引 5

用聚类音素表示实现零样本歌声转换,分离音色与演唱风格。

Zero-Shot Sing Voice Conversion: built upon clustering-based phoneme representations

  • 基于聚类的音素表征分离内容、音色与演唱风格。
  • 在超1万小时数据上测试,音质和音色准确率显著提升。
  • 适合需要零样本转换且关注韵律保留的研究者。

本研究提出一种创新的零样本任意对任意歌声转换(SVC)方法,利用新型聚类基音素表示,有效分离内容、音色与演唱风格,实现精准声线调控。研究发现,每位歌手录音较少的数据集更易出现音色泄露。在超过10,000小时的歌唱数据及用户反馈测试中,模型显著提升音质与音色准确性,达成预期目标,推动了语音转换技术发展。此外,该研究推进了零样本SVC的发展,为离散语音表示的未来工作奠定基础,尤其强调韵律的保持。

原文摘要 · Abstract (English)

This study presents an innovative Zero-Shot any-to-any Singing Voice Conversion (SVC) method, leveraging a novel clustering-based phoneme representation to effectively separate content, timbre, and singing style. This approach enables precise voice characteristic manipulation. We discovered that datasets with fewer recordings per artist are more susceptible to timbre leakage. Extensive testing on over 10,000 hours of singing and user feedback revealed our model significantly improves sound quality and timbre accuracy, aligning with our objectives and advancing voice conversion technology. Furthermore, this research advances zero-shot SVC and sets the stage for future work on discrete speech representation, emphasizing the preservation of rhyme.

歌声转换零样本音色分离聚类表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。