arXiv:2607.06929cs.SDcs.AI2026-07

构建9999首音乐的多维审美数据集,推动人机审美对齐研究

MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations

论文配图:MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations
图 1 · 摘自论文原文
  • 30名专业标注员对每首歌从10个维度评分,每首约10人标注
  • 现有模型预测与人类判断差距显著,暴露当前方法局限
  • 适合音乐理解、人机对齐、多模态分析等方向研究者使用

音乐审美评估是极具挑战但研究不足的问题,需模型捕捉精细、多维的人类感知判断。该领域进展受限于缺乏大规模、结构化审美标注数据。我们提出MADB,一个包含9,999首音乐作品的大规模数据集与基准,由30名训练过的标注员完成。每首音乐由约10位标注员在10个感知维度和1个总体评分上打分,并附有文本评论以支持多模态分析。我们在多个预训练模型上建立统一评估框架。结果揭示模型预测与人类判断间存在显著差距,暴露出当前方法的关键局限。MADB为人类对齐的音乐理解提供了新基准。

原文摘要 · Abstract (English)

Music aesthetic assessment is a challenging yet underexplored problem, requiring models to capture fine-grained, multi-dimensional human perceptual judgments. Progress in this area has been limited by the lack of large-scale datasets with structured aesthetic annotations. We introduce MADB, a large-scale dataset and benchmark comprising 9,999 tracks annotated by 30 trained annotators. Each track is rated by around 10 annotators across 10 perceptual dimensions and one overall score, with additional textual comments for multimodal analysis. We establish a unified evaluation framework over multiple pretrained models. Results reveal substantial gaps between model predictions and human judgments, exposing key limitations of current approaches. MADB provides a new benchmark for human-aligned music understanding. Project page: https://github.com/knownree/madb

音乐理解审美评估多维标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。