用人类反馈强化学习生成音乐,让系统随用户喜好自动优化。
Music Generation using Human-In-The-Loop Reinforcement Learning
- 结合音乐理论与人类反馈的强化学习框架
- 通过用户主观偏好实时提升生成音乐质量
- 适合音乐创作辅助与个性化作曲工具开发者
本文提出一种将人类在环强化学习(HITL RL)与音乐理论原则相结合的方法,实现音乐作品的实时生成。该方法借鉴了此前用于建模人形机器人动作和增强语言模型的应用经验,利用人类反馈优化训练过程。研究中构建了一个基于周期性表格Q-learning算法和epsilon-greedy探索策略的HITL RL框架,能够融合音乐理论的约束与规则。系统持续生成音乐片段,并通过迭代的人类反馈不断改进质量。整个过程的奖励函数由用户的主观音乐品味决定,使生成结果更符合个人审美。
原文摘要 · Abstract (English)
This paper presents an approach that combines Human-In-The-Loop Reinforcement Learning (HITL RL) with principles derived from music theory to facilitate real-time generation of musical compositions. HITL RL, previously employed in diverse applications such as modelling humanoid robot mechanics and enhancing language models, harnesses human feedback to refine the training process. In this study, we develop a HILT RL framework that can leverage the constraints and principles in music theory. In particular, we propose an episodic tabular Q-learning algorithm with an epsilon-greedy exploration policy. The system generates musical tracks (compositions), continuously enhancing its quality through iterative human-in-the-loop feedback. The reward function for this process is the subjective musical taste of the user.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。