用自学习方法提升音乐情感识别的准确率
Semi-Supervised Self-Learning Enhanced Music Emotion Recognition
- 通过分段训练+自学习机制减少标签噪声
- 在三个公开数据集上性能优于或相当现有方法
- 适合做音乐情感分析与小样本学习的研究者
音乐情感识别(MER)旨在识别音乐作品传达的情绪。然而,当前公开数据集样本量有限。近期提出基于片段的方法,将整个音频片段拆分为短段进行训练,自然扩充了样本且无需额外资源。随后将片段级预测结果聚合得到整首歌的判断。现有方法通常让片段继承其所在完整音频的标签,但音乐情绪在全段内并非恒定,这会引入标签噪声,导致过拟合。为此,本文提出一种半监督自学习(SSSL)方法,可在自学习过程中区分标签正确与错误的样本,从而有效利用扩增后的片段级数据。在三个公开情感数据集上的实验表明,该方法可取得更优或相当的性能。
原文摘要 · Abstract (English)
Music emotion recognition (MER) aims to identify the emotions conveyed in a given musical piece. However, currently, in the field of MER, the available public datasets have limited sample sizes. Recently, segment-based methods for emotion-related tasks have been proposed, which train backbone networks on shorter segments instead of entire audio clips, thereby naturally augmenting training samples without requiring additional resources. Then, the predicted segment-level results are aggregated to obtain the entire song prediction. The most commonly used method is that the segment inherits the label of the clip containing it, but music emotion is not constant during the whole clip. Doing so will introduce label noise and make the training easy to overfit. To handle the noisy label issue, we propose a semi-supervised self-learning (SSSL) method, which can differentiate between samples with correct and incorrect labels in a self-learning manner, thus effectively utilizing the augmented segment-level data. Experiments on three public emotional datasets demonstrate that the proposed method can achieve better or comparable performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。