提出分层增强的音乐审美评估框架,提升多维度音乐评价准确率
Hear: Hierarchically Enhanced Aesthetic Representations For Multidimensional Music Evaluation
- 分层融合片段与整首歌曲特征,捕捉多尺度音乐信息
- 在ICASSP 2026基准上各项指标均优于基线,显著提升评分精度
- 适合音乐推荐、智能创作等需精准审美判断的应用场景
由于音乐感知具有多维特性且标注数据稀缺,歌曲美学评价极具挑战。本文提出HEAR框架,包含:(1) 多源多尺度表征模块,获取互补的片段级与曲目级特征;(2) 分层增强策略以缓解过拟合;(3) 融合回归与排序损失的混合训练目标,实现精确打分与高阶歌曲可靠识别。实验表明,HEAR在ICASSP 2026 SongEval基准的两组测试中均持续优于基线,在所有指标上表现更优。代码与训练模型已开源。
原文摘要 · Abstract (English)
Evaluating song aesthetics is challenging due to the multidimensional nature of musical perception and the scarcity of labeled data. We propose HEAR, a robust music aesthetic evaluation framework that combines: (1) a multi-source multi-scale representations module to obtain complementary segment- and track-level features, (2) a hierarchical augmentation strategy to mitigate overfitting, and (3) a hybrid training objective that integrates regression and ranking losses for accurate scoring and reliable top-tier song identification. Experiments demonstrate that HEAR consistently outperforms the baseline across all metrics on both tracks of the ICASSP 2026 SongEval benchmark. The code and trained model weights are available at https://github.com/Eps-Acoustic-Revolution-Lab/EAR_HEAR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。