arXiv:2512.13998cs.SDcs.AI2025-12中稿 · manuscript被引 1

构建2496首乐曲情感数据集,提出自适应框架提升音乐情绪识别效果。

Memo2496: Expert-Annotated Dataset and Dual-view Adaptive Framework for Music Emotion Recognition

  • 双流注意力融合音谱图与耳蜗图,实现跨模态特征交互
  • 在三个数据集上达成最高情绪识别准确率,尤其在唤醒度上领先
  • 适合做音乐情感分析、人机交互与跨模态学习的研究者使用

音乐情绪识别(MER)受限于专家标注稀缺及跨异构数据集的鲁棒性不足。Memo2496 提供一个可复现的数据集,包含 2,496 首器乐作品的连续效价-唤醒度标签,由 30 名认证音乐专家标注,并通过界面熟悉与重复曲目内部标注者校准,在标准化圆形域中完成。本文还提出双视角自适应音乐情绪识别器(DAMER),在 Memo2496 和两个外部数据集上进行评估。DAMER 集成双流注意力融合(DSAF)实现梅尔谱图与耳蜗图间的词元级双向交互;渐进置信标签(PCL)通过温度调度与 Jensen-Shannon 散度生成课程式伪标签;风格锚定记忆学习(SAML)利用带标签的对比队列,正则化声学差异样本中的同情绪嵌入。主实验采用 PMEmo 与 1000songs 中常用的二分类 MER 协议,补充研究展示直接使用段级效价与唤醒度得分的回归能力。在 Memo2496、1000songs 与 PMEmo 上的实验表明,DAMER 在 Memo2496 与 1000songs 的唤醒度准确率上达到最优,在 PMEmo 的效价准确率上领先,且在 PMEmo 唤醒度上仍具竞争力。消融与诊断验证各模块有效性。数据集与源码已公开。

原文摘要 · Abstract (English)

Music Emotion Recognition (MER) is constrained by limited expert annotations and the need to establish robustness across heterogeneous corpora. Memo2496 supplies a reproducible dataset of 2,496 instrumental tracks with continuous valence-arousal labels from 30 certified music specialists, supported by interface familiarisation and duplicate-track intra-annotator calibration in a normalised circular domain. We also introduce the Dual-view Adaptive Music Emotion Recogniser (DAMER), a general framework evaluated on Memo2496 and two external datasets. DAMER integrates Dual-Stream Attention Fusion (DSAF) for token-level bidirectional interaction between Mel spectrograms and cochleagrams, Progressive Confidence Labelling (PCL) for curriculum-based pseudo-labels using temperature scheduling and Jensen-Shannon divergence, and Style-Anchored Memory Learning (SAML), whose labelled contrastive queue regularises same-emotion embeddings across acoustically varied samples. The primary evaluation follows the binary MER protocol used on PMEmo and 1000songs, while a supplementary continuous regression study demonstrates direct use of Memo2496 segment-level valence and arousal scores. Experiments on Memo2496, 1000songs, and PMEmo show that DAMER achieves the highest arousal accuracy among compared methods on Memo2496 and 1000songs and the highest valence accuracy on PMEmo, while remaining competitive for PMEmo arousal. Ablations and diagnostics validate each module. The dataset and source code are publicly available.

音乐情绪识别数据集双流网络自适应学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。