arXiv:2512.17946cs.SDcs.AI2025-12AAAI被引 1

通过注入调式信息,提升模型对音乐情感的识别能力。

Let the Model Learn to Feel: Mode-Guided Tonality Injection for Symbolic Music Emotion Recognition

  • 在预训练模型中引入调式特征,增强情感表征
  • 在EMOPIA和VGMIDI上分别达到75.2%和59.1%准确率
  • 适合关注音乐心理学与情感建模的研究者

音乐情感识别是符号化音乐理解(SMER)的关键任务。近期方法通过微调大规模预训练模型(如MIDIBERT,符号化音乐理解的基准模型)将音乐语义映射到情感标签,取得良好效果。然而,这些模型虽能捕捉音乐的分布语义,却常忽略调式结构——而调式在音乐心理学中对情感感知至关重要。本文研究了MIDIBERT在调式-情感关联上的表征能力,发现其存在局限。为此,提出模式引导增强(MoGE)策略,结合音乐心理学先验知识。首先通过调式增强分析揭示MIDIBERT未能有效编码情绪-调式关联;随后定位最不相关的情感层,并设计一种调式引导的逐特征线性调制注入(MoFi)框架,显式注入调式特征,以提升情感表示与推理能力。在EMOPIA和VGMIDI数据集上的大量实验表明,该策略显著提升SMER性能,准确率分别达75.2%和59.1%,验证了调式引导建模在符号化音乐情感识别中的有效性。

原文摘要 · Abstract (English)

Music emotion recognition is a key task in symbolic music understanding (SMER). Recent approaches have shown promising results by fine-tuning large-scale pre-trained models (e.g., MIDIBERT, a benchmark in symbolic music understanding) to map musical semantics to emotional labels. While these models effectively capture distributional musical semantics, they often overlook tonal structures, particularly musical modes, which play a critical role in emotional perception according to music psychology. In this paper, we investigate the representational capacity of MIDIBERT and identify its limitations in capturing mode-emotion associations. To address this issue, we propose a Mode-Guided Enhancement (MoGE) strategy that incorporates psychological insights on mode into the model. Specifically, we first conduct a mode augmentation analysis, which reveals that MIDIBERT fails to effectively encode emotion-mode correlations. We then identify the least emotion-relevant layer within MIDIBERT and introduce a Mode-guided Feature-wise linear modulation injection (MoFi) framework to inject explicit mode features, thereby enhancing the model's capability in emotional representation and inference. Extensive experiments on the EMOPIA and VGMIDI datasets demonstrate that our mode injection strategy significantly improves SMER performance, achieving accuracies of 75.2% and 59.1%, respectively. These results validate the effectiveness of mode-guided modeling in symbolic music emotion recognition.

音乐情感调式注入MIDIBERT符号化音乐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。