arXiv:2512.10264cs.SD2025-12被引 4

用多重奖励优化音乐生成,让模型更懂人类喜好。

MR-FlowDPO: Multi-Reward Direct Preference Optimization for Flow-Matching Text-to-Music Generation

  • 引入多维度奖励机制,评估文本对齐、音质和语义一致性。
  • 生成音乐在音质、对齐度和韵律稳定性上显著优于基线。
  • 适合关注音乐生成质量与人机对齐的研究者与开发者。

音乐生成模型面临难以直接对齐人类偏好的挑战,因音乐评价具有高度主观性且个体差异大。本文提出MR-FlowDPO,一种基于流匹配的音乐生成模型优化方法,通过多奖励直接偏好优化(DPO)提升生成质量。该方法设计三个维度的奖励:文本对齐、音频制作质量与语义一致性,均采用可扩展的现成模型进行预测。奖励用于构建偏好数据集或融入文本提示。针对音乐性评估模糊的问题,提出基于语义自监督表示的新评分机制,显著提升生成音乐的节奏稳定性。通过多种音乐专用客观指标与人工评测验证,结果表明MR-FlowDPO在音质、文本对齐与音乐性方面均显著优于多个先进基线。代码与演示地址公开于https://github.com/lonzi/mrflow_dpo及https://lonzi.github.io/mr_flowdpo_demopage/。

原文摘要 · Abstract (English)

A key challenge in music generation models is their lack of direct alignment with human preferences, as music evaluation is inherently subjective and varies widely across individuals. We introduce MR-FlowDPO, a novel approach that enhances flow-matching-based music generation models - a major class of modern music generative models, using Direct Preference Optimization (DPO) with multiple musical rewards. The rewards are crafted to assess music quality across three key dimensions: text alignment, audio production quality, and semantic consistency, utilizing scalable off-the-shelf models for each reward prediction. We employ these rewards in two ways: (i) By constructing preference data for DPO and (ii) by integrating the rewards into text prompting. To address the ambiguity in musicality evaluation, we propose a novel scoring mechanism leveraging semantic self-supervised representations, which significantly improves the rhythmic stability of generated music. We conduct an extensive evaluation using a variety of music-specific objective metrics as well as a human study. Results show that MR-FlowDPO significantly enhances overall music generation quality and is consistently preferred over highly competitive baselines in terms of audio quality, text alignment, and musicality. Our code is publicly available at https://github.com/lonzi/mrflow_dpo. Samples are provided in our demo page at https://lonzi.github.io/mr_flowdpo_demopage/.

音乐生成偏好优化流匹配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。