arXiv:2506.19085cs.LGcs.SD2025-06中稿 · ICASSP 2025被引 22

通过真人对比测试,首次对主流音乐生成模型和评估指标进行真实偏好排名。

Benchmarking Music Generation Models and Metrics via Human Preference Studies

  • 用12个顶尖模型生成6000首歌,让2500人做1.5万次两两对比
  • 发现现有指标与人类偏好相关性普遍较低,部分指标甚至反向误导
  • 开源全部数据,推动音乐生成评价的主观化研究

近期进展使生成音乐日益接近人类创作,但模型评估仍具挑战。尽管人类偏好是质量评估的金标准,但将其转化为客观指标(尤其是文本-音频对齐与音乐质量)仍困难重重。本文使用12个前沿模型生成6000首歌曲,通过2500名参与者完成1.5万次成对音频比较,评估人类偏好与常用指标的相关性。据我们所知,这是首个基于人类偏好的主流音乐生成模型与评估指标排名工作。为推进主观评价研究,本文公开发布生成音乐数据集与人类评估结果。

原文摘要 · Abstract (English)

Recent advancements have brought generated music closer to human-created compositions, yet evaluating these models remains challenging. While human preference is the gold standard for assessing quality, translating these subjective judgments into objective metrics, particularly for text-audio alignment and music quality, has proven difficult. In this work, we generate 6k songs using 12 state-of-the-art models and conduct a survey of 15k pairwise audio comparisons with 2.5k human participants to evaluate the correlation between human preferences and widely used metrics. To the best of our knowledge, this work is the first to rank current state-of-the-art music generation models and metrics based on human preference. To further the field of subjective metric evaluation, we provide open access to our dataset of generated music and human evaluations.

音乐生成人类偏好评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。