首个大规模歌曲美学评估数据集,助力生成音乐更符合人类审美。
SongEval: A Benchmark Dataset for Song Aesthetics Evaluation
- 构建包含2399首全曲的跨语言歌曲美学数据集
- 16位专业音乐人从5个维度评分,总时长超140小时
- 比现有客观指标更能预测人类对音乐质量的感知
美学是歌曲生成任务中反映人类感知的重要隐性标准,但生成歌曲的美学评估仍具挑战,因音乐欣赏高度主观。现有基于嵌入距离的评价指标难以捕捉主观与感知层面的音乐吸引力。为此,我们提出SongEval,首个开源、大规模的完整歌曲美学评估基准数据集。SongEval包含2399首全曲,总计超过140小时,由16位具有音乐背景的专业标注者进行评分。每首歌在五个关键维度上评估:整体连贯性、记忆度、人声呼吸与乐句自然性、歌曲结构清晰度及整体音乐性。数据集涵盖英文与中文歌曲,覆盖九种主流音乐流派。为验证评估有效性,我们利用SongEval预测美学得分,结果表明其性能优于现有客观评价指标,在预测人类感知音乐质量方面表现更佳。
原文摘要 · Abstract (English)
Aesthetics serve as an implicit and important criterion in song generation tasks that reflect human perception beyond objective metrics. However, evaluating the aesthetics of generated songs remains a fundamental challenge, as the appreciation of music is highly subjective. Existing evaluation metrics, such as embedding-based distances, are limited in reflecting the subjective and perceptual aspects that define musical appeal. To address this issue, we introduce SongEval, the first open-source, large-scale benchmark dataset for evaluating the aesthetics of full-length songs. SongEval includes over 2,399 songs in full length, summing up to more than 140 hours, with aesthetic ratings from 16 professional annotators with musical backgrounds. Each song is evaluated across five key dimensions: overall coherence, memorability, naturalness of vocal breathing and phrasing, clarity of song structure, and overall musicality. The dataset covers both English and Chinese songs, spanning nine mainstream genres. Moreover, to assess the effectiveness of song aesthetic evaluation, we conduct experiments using SongEval to predict aesthetic scores and demonstrate better performance than existing objective evaluation metrics in predicting human-perceived musical quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。