JAM实现歌词级精准控制,生成更符合人类审美的小型化音乐模型。
JAM: A Tiny Flow-based Song Generator with Fine-grained Controllability and Aesthetic Alignment
- 基于流匹配的端到端音乐生成,支持词级时序与持续时间控制。
- 通过直接偏好优化提升音质,使生成歌曲更贴近人类审美。
- 提供公开评估数据集JAME,推动歌词到歌曲模型标准化评测。
扩散模型和流匹配模型近年来彻底改变了自动文本到音频的生成。这些模型能生成高质量且忠实于语音与声学事件的音频输出。然而,在以音乐和歌曲为主的创造性音频生成方面仍有改进空间。近期的开源歌词到歌曲模型如DiffRhythm、ACE-Step和LeVo已达到娱乐用途的可接受标准,但缺乏音乐人工作流中常见的细粒度词级可控性。据我们所知,基于流匹配的JAM是首个实现歌词级时序与持续时间控制的模型,支持精细的人声调控。为提升生成歌曲质量以更好地契合人类偏好,我们采用直接偏好优化(DPO)进行美学对齐,通过合成数据集迭代优化模型,无需人工标注。此外,我们推出了公开评估数据集JAME,旨在统一此类歌词到歌曲模型的评价标准。实验表明,JAM在音乐特定属性上优于现有模型。
原文摘要 · Abstract (English)
Diffusion and flow-matching models have revolutionized automatic text-to-audio generation in recent times. These models are increasingly capable of generating high quality and faithful audio outputs capturing to speech and acoustic events. However, there is still much room for improvement in creative audio generation that primarily involves music and songs. Recent open lyrics-to-song models, such as, DiffRhythm, ACE-Step, and LeVo, have set an acceptable standard in automatic song generation for recreational use. However, these models lack fine-grained word-level controllability often desired by musicians in their workflows. To the best of our knowledge, our flow-matching-based JAM is the first effort toward endowing word-level timing and duration control in song generation, allowing fine-grained vocal control. To enhance the quality of generated songs to better align with human preferences, we implement aesthetic alignment through Direct Preference Optimization, which iteratively refines the model using a synthetic dataset, eliminating the need or manual data annotations. Furthermore, we aim to standardize the evaluation of such lyrics-to-song models through our public evaluation dataset JAME. We show that JAM outperforms the existing models in terms of the music-specific attributes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。