音乐评分模型常因流派偏见误判质量,新方法让模型更关注音乐本身而非类型。
Genre Bias or Aesthetic Perception? Identifying and Mitigating Shortcut Learning in Music Evaluation

- 通过重加权难样本并约束群体表现,减少流派相关偏见。
- 在跨流派和同流派偏好上均提升与人类判断的一致性。
- 适合用于生成模型评估、数据集筛选等需公平评分的场景。
音乐审美评分在数据集构建、生成模型评估及音乐生成奖励建模中至关重要。现有方法依赖人工标注评分训练深度神经网络,但模型可能利用虚假相关性而非真实美学特征。本文通过系统分析SongEval发现:训练数据中的流派偏差导致流派特征与预测分数强相关,使模型以流派为捷径判断质量,造成流行音乐被系统性高估,其他高质量作品被低估,结果与人类偏好不一致。为此,我们提出一种联合重加权难样本与正则化群体性能的训练目标,促使模型学习与流派无关的音乐性表征。实验表明,该方法显著降低流派依赖偏差,并提升与人类偏好的一致性,跨流派与同流派对齐均有改善。
原文摘要 · Abstract (English)
Music aesthetics scoring plays a critical role in applications such as dataset curation, generative model evaluation, and reward modeling for music generation. Recent approaches rely on deep neural networks trained on human-annotated ratings, but these models may exploit spurious correlations rather than capturing perceptually meaningful aesthetics. In this work, we identify a previously underexplored failure mode in music evaluation models: genre-induced shortcut learning. Through a systematic analysis of SongEval, we show that biases in training data lead to strong correlations between genre-related features and predicted scores, causing the model to use them as a proxy for aesthetics. This results in systematic overestimation of pop music and undervaluation of high-quality samples from other genres, leading to predictions that are inconsistent with human preferences. To address this issue, we propose a training objective that jointly reweights hard samples and regularizes group-level performance, encouraging the model to learn genre-invariant representations of musicality. Experimental results demonstrate that our method reduces genre-dependent bias and improves alignment with human preferences, as reflected by gains in both cross-genre and within-genre preference alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。