跨数据集音乐情感识别存在分布差距,提出融合音高特征与模型嵌入的改进方案。
A Study on the Data Distribution Gap in Music Emotion Recognition
- 整合五大数据集,分析不同音乐风格下情感标注的分布差异
- 发现特定风格在特征表示中存在主导现象,影响模型泛化能力
- 结合Jukebox嵌入与音高特征,显著提升跨数据集识别效果
音乐情感识别(MER)依赖于人类主观标注,以往研究多聚焦单一音乐风格,缺乏对摇滚、古典等多元风格的统一建模。本文系统考察了包含维度情感标注的五个数据集——EmoMusic、DEAM、PMEmo、WTC和WCMED,涵盖多种音乐风格。通过实验揭示了模型在分布外数据上泛化能力差的问题。深入分析多个数据与特征集后,发现现有数据中存在风格-情感关联模式,并揭示某些特征表示中的风格主导与数据集偏差。基于此,提出一种简单有效的框架:将Jukebox模型提取的嵌入与音高特征(chroma features)结合,并利用多个多样化训练集进行训练,显著提升了模型在跨数据集场景下的泛化性能。
原文摘要 · Abstract (English)
Music Emotion Recognition (MER) is a task deeply connected to human perception, relying heavily on subjective annotations collected from contributors. Prior studies tend to focus on specific musical styles rather than incorporating a diverse range of genres, such as rock and classical, within a single framework. In this paper, we address the task of recognizing emotion from audio content by investigating five datasets with dimensional emotion annotations -- EmoMusic, DEAM, PMEmo, WTC, and WCMED -- which span various musical styles. We demonstrate the problem of out-of-distribution generalization in a systematic experiment. By closely looking at multiple data and feature sets, we provide insight into genre-emotion relationships in existing data and examine potential genre dominance and dataset biases in certain feature representations. Based on these experiments, we arrive at a simple yet effective framework that combines embeddings extracted from the Jukebox model with chroma features and demonstrate how, alongside a combination of several diverse training sets, this permits us to train models with substantially improved cross-dataset generalization capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。