首个专家评分的文本生成音乐数据集,助力自动评估更贴近人类听感。
MusicEval: A Generative Music Dataset with Expert Ratings for Automatic Text-to-Music Evaluation
- 构建首个基于专家评分的文本到音乐生成评估数据集
- 包含2748段生成音乐与13740条专家评分,覆盖31个主流模型
- 适用于音乐生成、评估算法研究者及艺术科技交叉领域
文本到音乐(TTM)生成技术发展迅速,但评估仍面临挑战,因现有客观与主观方法难以兼顾效果与成本。本文提出一种对齐人类感知的自动化评估任务,并构建MusicEval——首个生成式音乐评估数据集。该数据集包含31个先进模型基于384个文本提示生成的2748段音乐片段,以及14位音乐专家提供的13740条评分。基于此数据集,设计了基于CLAP的评估模型,实验验证了该任务的可行性,为未来TTM评估研究提供了重要参考。数据集已公开:https://www.aishelltech.com/AISHELL_7A。
原文摘要 · Abstract (English)
The technology for generating music from textual descriptions has seen rapid advancements. However, evaluating text-to-music (TTM) systems remains a significant challenge, primarily due to the difficulty of balancing performance and cost with existing objective and subjective evaluation methods. In this paper, we propose an automatic assessment task for TTM models to align with human perception. To address the TTM evaluation challenges posed by the professional requirements of music evaluation and the complexity of the relationship between text and music, we collect MusicEval, the first generative music assessment dataset. This dataset contains 2,748 music clips generated by 31 advanced and widely used models in response to 384 text prompts, along with 13,740 ratings from 14 music experts. Furthermore, we design a CLAP-based assessment model built on this dataset, and our experimental results validate the feasibility of the proposed task, providing a valuable reference for future development in TTM evaluation. The dataset is available at https://www.aishelltech.com/AISHELL_7A.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。