统一处理音乐情感的分类与维度标签,提升跨数据集识别效果。
Towards Unified Music Emotion Recognition across Dimensional and Categorical Models
- 融合类别与维度情感标签,构建多任务学习框架。
- 在多个数据集上优于当前最优模型,显著提升泛化能力。
- 适合需要跨模态情感分析的研究者和开发者。
音乐情感识别(MER)面临的主要挑战之一是不同数据集采用异构的情感表示方式,如类别标签(如快乐、悲伤)与维度标签(如愉悦度-唤醒度)。本文提出一种统一的多任务学习框架,可同时处理两类标签,并在多个数据集上联合训练。该框架结合音乐特征(调性、和弦)与MERT嵌入作为输入表示,并采用知识蒸馏技术,将各独立数据集上训练的教师模型知识迁移至学生模型,增强其跨任务泛化能力。在MTG-Jamendo、DEAM、PMEmo和EmoMusic等多个数据集上的实验表明,引入音乐特征、多任务学习和知识蒸馏显著提升性能。特别地,本模型在MTG-Jamendo数据集上超越了MediaEval 2021竞赛最佳模型的表现。本工作为MER提供了统一整合类别与维度标签的框架,支持跨数据集训练。
原文摘要 · Abstract (English)
One of the most significant challenges in Music Emotion Recognition (MER) comes from the fact that emotion labels can be heterogeneous across datasets with regard to the emotion representation, including categorical (e.g., happy, sad) versus dimensional labels (e.g., valence-arousal). In this paper, we present a unified multitask learning framework that combines these two types of labels and is thus able to be trained on multiple datasets. This framework uses an effective input representation that combines musical features (i.e., key and chords) and MERT embeddings. Moreover, knowledge distillation is employed to transfer the knowledge of teacher models trained on individual datasets to a student model, enhancing its ability to generalize across multiple tasks. To validate our proposed framework, we conducted extensive experiments on a variety of datasets, including MTG-Jamendo, DEAM, PMEmo, and EmoMusic. According to our experimental results, the inclusion of musical features, multitask learning, and knowledge distillation significantly enhances performance. In particular, our model outperforms the state-of-the-art models, including the best-performing model from the MediaEval 2021 competition on the MTG-Jamendo dataset. Our work makes a significant contribution to MER by allowing the combination of categorical and dimensional emotion labels in one unified framework, thus enabling training across datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。