arXiv:2508.05011cs.SDcs.AI2025-08被引 1

用强化学习优化偏好,让生成的歌不跑题、更贴合歌词。

Towards Hallucination-Free Music: A Reinforcement Learning Preference Optimization Framework for Reliable Song Generation

  • 基于语音错误率构建人类偏好数据集,精准捕捉歌词对齐需求。
  • DPO方法使语音错误率降低7.4%,显著减少内容幻觉。
  • 框架可迁移,适合追求歌词忠实度与音乐质量的研究者。

近年来,基于音频的生成语言模型加速了从歌词生成歌曲的进程。然而,这些模型常出现内容幻觉,导致输出与输入歌词不符,破坏音乐连贯性。当前的监督微调(SFT)方法受限于被动标签拟合,自我改进能力弱且难以缓解幻觉问题。为此,我们提出一种新颖的强化学习(RL)框架,利用偏好优化实现幻觉控制。关键贡献包括:(1) 构建基于音素错误率(PER)计算和规则过滤的鲁棒幻觉偏好数据集,以捕捉人类期望的对齐;(2) 在RL框架中实施并评估三种不同偏好优化策略:直接偏好优化(DPO)、近端策略优化(PPO)和组相对策略优化(GRPO)。DPO采用离策略,提升正向词元概率,实现7.4%的PER下降。PPO和GRPO采用同策略,训练基于PER的奖励模型,通过奖励最大化与KL正则化迭代优化序列,分别带来4.9%和4.7%的PER下降。全面的客观与主观评估表明,该方法有效抑制幻觉,同时保持音乐质量。本工作提供了一套系统化的基于强化学习的幻觉控制方案,其可迁移性也为音乐风格遵循与音乐性提升开辟新路径。

原文摘要 · Abstract (English)

Recent advances in audio-based generative language models have accelerated AI-driven lyric-to-song generation. However, these models frequently suffer from content hallucination, producing outputs misaligned with the input lyrics and undermining musical coherence. Current supervised fine-tuning (SFT) approaches, limited by passive label-fitting, exhibit constrained self-improvement and poor hallucination mitigation. To address this core challenge, we propose a novel reinforcement learning (RL) framework leveraging preference optimization for hallucination control. Our key contributions include: (1) Developing a robust hallucination preference dataset constructed via phoneme error rate (PER) computation and rule-based filtering to capture alignment with human expectations; (2) Implementing and evaluating three distinct preference optimization strategies within the RL framework: Direct Preference Optimization (DPO), Proximal Policy Optimization (PPO), and Group Relative Policy Optimization (GRPO). DPO operates off-policy to enhance positive token likelihood, achieving a significant 7.4% PER reduction. PPO and GRPO employ an on-policy approach, training a PER-based reward model to iteratively optimize sequences via reward maximization and KL-regularization, yielding PER reductions of 4.9% and 4.7%, respectively. Comprehensive objective and subjective evaluations confirm that our methods effectively suppress hallucinations while preserving musical quality. Crucially, this work presents a systematic, RL-based solution to hallucination control in lyric-to-song generation. The framework's transferability also unlocks potential for music style adherence and musicality enhancement, opening new avenues for future generative song research.

音乐生成强化学习幻觉抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。