arXiv:2608.30559cs.SD2026-08

用空间启发奖励训练音频语言模型,自动实现音乐混音升维。

SPHERE: Automatic Music Upmixing via Audio Language Model Post-Training with Spatial Heuristic Rewards

论文配图:SPHERE: Automatic Music Upmixing via Audio Language Model Post-Training with Spatial Heuristic Rewards
图 1 · 摘自论文原文
  • 基于音频语言模型,通过拒绝采样和强化学习优化混音参数。
  • 使用6个感知驱动的奖励项,使混音更中心、平衡且有空间感。
  • 无需专用架构,可将专家知识转化为可验证奖励注入模型。

本文研究自动音乐升维任务,即从多音轨录音中预测空间混音参数。不同于依赖任务特定音乐编码器的现有方法,我们采用音频语言模型(ALM)后训练策略,利用已有ALM中蕴含的丰富音乐语义与混音知识。具体提出一种后训练方案:先进行拒绝采样监督微调(SFT),再通过GRPO算法实施基于可验证奖励的强化学习(RLVR)。我们设计了名为Sphere(空间启发奖励)的确定性奖励体系,其灵感源自音乐混音规范,包含6个感知动机子奖励,引导输出混音具备中心化、均衡性与空间感。实验表明,专家领域知识可通过可验证奖励形式注入语言模型,而无需定制化架构。

原文摘要 · Abstract (English)

In this paper, we study the task of automatic music upmixing, wherein a system predicts spatial mixing parameters from a multi-stem recording. Different from existing methods that rely on task-specific music encoders, we approach this task via audio language model (ALM) post-training, leveraging rich representations from existing ALMs, which encode both music semantics and mixing knowledge. Specifically, we propose a post-training recipe that first employs rejection sampling SFT, followed by reinforcement learning (RL) with verifiable rewards (RLVR) via GRPO. We propose Sphere (Spatial Heuristic Rewards), a deterministic reward suite inspired by music mixing conventions, to guide our post-training. It consists of 6 perceptually-motivated sub-rewards and encourages the output mix to be centered, balanced and spacious. More broadly, our results suggest that expert domain knowledge can be encoded as verifiable rewards and distilled into language models, without task-specific architectures.

音乐生成音频语言模型强化学习混音升维

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。