让歌声更个性化,支持多种歌唱任务的通用音高生成模型
StylePitcher: Generating Style-Following and Expressive Pitch Curves for Versatile Singing Tasks
- 基于修正流匹配架构,从参考音频学习歌手风格
- 在保持音准准确的同时,提升风格相似度与音频质量
- 无需重训练即可适配音高矫正、合成人声等多样任务
现有音高曲线生成方法面临两大挑战:通常忽略歌手特有的表现力,难以捕捉个体演唱风格;且多作为特定任务(如音高校正、合成语音或语音转换)的辅助模块,泛化能力受限。我们提出 StylePitcher,一种通用音高曲线生成器,能从参考音频中学习歌手风格,同时保持与目标旋律的对齐。该模型基于修正流匹配架构,灵活引入乐谱符号和音高上下文作为条件,可无缝适应多种歌唱任务而无需重新训练。跨多种歌唱任务的客观与主观评估表明,StylePitcher 在提升风格相似度和音频质量的同时,保持与专用基线相当的音高准确性。
原文摘要 · Abstract (English)
Existing pitch curve generators face two main challenges: they often neglect singer-specific expressiveness, reducing their ability to capture individual singing styles. And they are typically developed as auxiliary modules for specific tasks such as pitch correction, singing voice synthesis, or voice conversion, which restricts their generalization capability. We propose StylePitcher, a general-purpose pitch curve generator that learns singer style from reference audio while preserving alignment with the intended melody. Built upon a rectified flow matching architecture, StylePitcher flexibly incorporates symbolic music scores and pitch context as conditions for generation, and can seamlessly adapt to diverse singing tasks without retraining. Objective and subjective evaluations across various singing tasks demonstrate that StylePitcher improves style similarity and audio quality while maintaining pitch accuracy comparable to task-specific baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。