Muskits-ESPnet用预训练音频模型革新歌唱合成,支持多格式输入与自动评估。
Muskits-ESPnet: A Comprehensive Toolkit for Singing Voice Synthesis in New Paradigm
- 融合预训练音频模型,支持连续与离散表征的歌唱合成新范式。
- 可自动检测修正乐谱错误,并模拟人类主观评分进行质量评估。
- 适合语音合成、音乐生成研究者,尤其关注智能流程与多模态输入。
本研究提出Muskits-ESPnet,一个多功能工具包,通过在连续与离散两种路径中应用预训练音频模型,为歌唱语音合成(SVS)引入新范式。具体而言,探索了来自自监督学习(SSL)模型和音频编解码器的离散表示,显著提升系统的灵活性与智能化水平,支持多格式输入及可适配的数据处理流程,适用于多种SVS模型。该工具包包含自动音乐乐谱错误检测与修正功能,以及感知自动评估模块,可模拟人类主观评分。Muskits-ESPnet开源地址:https://github.com/espnet/espnet。
原文摘要 · Abstract (English)
This research presents Muskits-ESPnet, a versatile toolkit that introduces new paradigms to Singing Voice Synthesis (SVS) through the application of pretrained audio models in both continuous and discrete approaches. Specifically, we explore discrete representations derived from SSL models and audio codecs and offer significant advantages in versatility and intelligence, supporting multi-format inputs and adaptable data processing workflows for various SVS models. The toolkit features automatic music score error detection and correction, as well as a perception auto-evaluation module to imitate human subjective evaluating scores. Muskits-ESPnet is available at \url{https://github.com/espnet/espnet}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。