arXiv:2603.00610cs.SDcs.AI2026-03中稿 · ICML被引 5

为多模态音乐生成设计新评估体系,提升模型与人类判断的一致性。

CMI-RewardBench: Evaluating Music Reward Models with Compositional Multimodal Instruction

  • 构建基于组合式多模态指令的音乐奖励评估框架
  • 在11万条伪标注数据上训练出与人类评分高度相关的奖励模型
  • 支持文本、歌词、音频混合指令,适合音乐生成研究者使用

尽管音乐生成模型已能处理包含文本、歌词和参考音频的复杂多模态输入,但评估机制仍滞后。本文通过建立面向组合式多模态指令(CMI)的音乐奖励建模完整生态,填补这一空白。提出CMI-Pref-Pseudo,一个包含11万条伪标注样本的大规模偏好数据集,以及专为细粒度对齐任务设计的高质量人工标注语料CMI-Pref。为统一评估标准,构建CMI-RewardBench,该基准在音乐性、文本-音乐对齐、组合指令对齐三个维度上评估音乐奖励模型。基于这些资源,开发参数高效型奖励模型家族CMI-RMs,可处理异构输入。实验表明,CMI-RM在音乐性和对齐性上与人类评分高度相关,并可通过top-k过滤实现推理阶段的有效扩展。代码、模型权重及数据集均已开源。

原文摘要 · Abstract (English)

While music generation models have evolved to handle complex multimodal inputs mixing text, lyrics, and reference audio, evaluation mechanisms have lagged behind. In this paper, we bridge this critical gap by establishing a comprehensive ecosystem for music reward modeling under Compositional Multimodal Instruction (CMI), where the generated music may be conditioned on text descriptions, lyrics, and audio prompts. We first introduce CMI-Pref-Pseudo, a large-scale preference dataset comprising 110k pseudo-labeled samples, and CMI-Pref, a high-quality, human-annotated corpus tailored for fine-grained alignment tasks. To unify the evaluation landscape, we propose CMI-RewardBench, a unified benchmark that evaluates music reward models on heterogeneous samples across musicality, text-music alignment, and compositional instruction alignment. Leveraging these resources, we develop CMI reward models (CMI-RMs), a parameter-efficient reward model family capable of processing heterogeneous inputs. We evaluate their correlation with human judgment scores on musicality and alignment on CMI-Pref along with previous datasets. Further experiments demonstrate that CMI-RM not only correlates strongly with human judgments, but also enables effective inference-time scaling via top-k filtering. Code is available at GitHub (https://github.com/Haiwen-Xia/CMI-RewardBench). Model weights: CMI-RM (https://huggingface.co/HaiwenXia/CMI-RM). Datasets: CMI-Pref-Pseudo (https://huggingface.co/datasets/HaiwenXia/cmi-pref-pseudo) and CMI-Pref (https://huggingface.co/datasets/HaiwenXia/cmi-pref)

音乐生成多模态评估奖励模型人机对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。