用规则约束提升歌词生成旋律的音乐合理性
Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints

- 基于音乐规则自动生成偏好数据,无须人工标注
- 经两阶段优化后,规则违反率显著降低,旋律更连贯
- 适合音乐生成、人机共创等场景的研究者与创作者
大型语言模型在歌词到旋律生成中展现潜力,但经监督微调(SFT)训练的模型常生成不合理的旋律,如节奏不当、人声音域不合适,我们称之为“规则违背”。为解决此问题,提出一种无需人工标注的对齐框架:通过定义规则化音乐约束,自动从SFT模型输出中构建偏好数据集。模型采用分步优化策略,先用直接偏好优化(DPO)处理成对偏好数据,再用Kahneman-Tversky优化(KTO)处理非配对负样本。实验表明,该对齐模型显著减少规则违背,在客观与主观评估中均优于强基线,生成旋律的音乐性与连贯性大幅提升。交互式演示及音频对比可访问 https://arain233.github.io/AligningMelody-demo。
原文摘要 · Abstract (English)
Large Language Models (LLMs) show promise in lyric-to-melody generation, but models trained with Supervised Fine-Tuning (SFT) often produce musically implausible melodies with issues like poor rhythm and unsuitable vocal ranges, a phenomenon we term "constraint violation". To address this, we propose a novel alignment framework that instills musical knowledge without human annotation. We define rule-based musical constraints to automatically generate a preference dataset from an SFT model's outputs. The model is then aligned through a sequential process, first using Direct Preference Optimization (DPO) on paired preference data, followed by Kahneman-Tversky Optimization (KTO) on unpaired negative samples. Experimental results demonstrate that our aligned model substantially reduces rule violations and outperforms strong baselines in both objective and subjective evaluations, generating melodies with substantially improved musicality and coherence. An interactive demo with audio comparisons is available at https://arain233.github.io/AligningMelody-demo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。