arXiv:2504.16839cs.SD2025-04被引 1

用音频审美评分优化音乐生成,提升听感但需防过度单调

SMART: Tuning a symbolic music generation system with an audio domain aesthetic reward

  • 用音频审美模型作为奖励,通过强化学习微调钢琴音乐生成
  • 初步听感测试中14人评分平均提升,低层特征也发生改变
  • 过度优化会严重降低生成多样性,适合注重质量的创作者

近期研究提出训练机器学习模型以预测音乐音频的审美评分。本文探索此类模型能否用于强化学习微调符号化音乐生成系统,并分析其对输出的影响。为此,我们采用群体相对策略优化方法,以Meta Audiobox美学评分作为音频渲染输出的奖励信号,微调一个钢琴MIDI生成模型。实验发现,该优化过程显著影响生成结果的多个低层特征,并在包含14名参与者的初步听感测试中提升了平均主观评分。同时,过度优化会导致模型输出多样性急剧下降。

原文摘要 · Abstract (English)

Recent work has proposed training machine learning models to predict aesthetic ratings for music audio. Our work explores whether such models can be used to finetune a symbolic music generation system with reinforcement learning, and what effect this has on the system outputs. To test this, we use group relative policy optimization to finetune a piano MIDI model with Meta Audiobox Aesthetics ratings of audio-rendered outputs as the reward. We find that this optimization has effects on multiple low-level features of the generated outputs, and improves the average subjective ratings in a preliminary listening study with $14$ participants. We also find that over-optimization dramatically reduces diversity of model outputs.

音乐生成强化学习审美评估多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。