arXiv:2606.21670cs.SDcs.AI2026-06被引 1

用人类偏好奖励提升文本生成音乐质量,效果显著。

Improving Text-to-Music Generation with Human Preference Rewards

  • 引入人类偏好奖励作为训练与推理的双重信号。
  • 专家迭代使性能提升最明显,达基准的2.1倍。
  • 适合关注音乐生成质量与人类感知对齐的研究者。

我们介绍在 ICME 2026 学术文本转音乐(ATTM)挑战赛效率赛道中的参赛方案。除了挑战赛标准的 FAD-CLAP 与 CLAP 分数外,我们引入了来自 TuneJury 的学习型人类偏好奖励,该奖励由在开放音乐偏好数据集上训练的双塔排序器获得。该奖励既用于训练时的条件输入,也用于采样选择。整个流程基于 120M 参数的 FluxAudio-S 骨干网络,包含五项工程决策:(i) 训练时奖励条件化,同时作为推理时的类条件生成(CFG)轴;(ii) 对五种得分条件化架构进行搜索,训练与推理使用不同变体;(iii) 对前 10% 优秀样本进行专家迭代;(iv) 通过轻量级偏好微调(CRPO)提升音频-文本对齐;(v) 推理后处理采用联合 CFG、源分离与响度归一化。在 100 个 Song Describer 提示上的逐阶段分析显示:训练时奖励条件化为有效条件轴,专家迭代贡献最大,偏好微调仅带来噪声级别提升,而推理时的得分标量在链路末端已趋于饱和。

原文摘要 · Abstract (English)

We describe our entry to the efficiency track of the Academic Text-to-Music (ATTM) Grand Challenge at ICME 2026. Beyond the challenge protocol's FAD-CLAP and CLAP score, we add a learned human-preference reward from TuneJury, a twin pairwise ranker trained over open music-preference datasets. The reward serves both as a training-time conditioning signal and as a sample-selection criterion. The pipeline combines five engineering decisions on a 120M-parameter FluxAudio-S backbone, four at training time and one at inference: (i) training-time reward conditioning that doubles as an inference-time CFG axis, (ii) a sweep over five score-conditioning architectures, where training and inference use different variants, (iii) expert iteration on the top decile, (iv) a short preference-tuning pass (CRPO) for audio-text alignment, and (v) inference post-processing via joint CFG, source separation, and loudness normalization. Per-stage decomposition on 100 Song Describer prompts shows training-time reward conditioning as a functional conditioning axis, expert iteration as the dominant contributor, the preference-tuning pass adding only noise-level gain, and the inference-time score scalar already saturated by the end of the chain.

文本生成音乐人类偏好生成优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。