基于双尺度注意力的元学习模型,实现个性化音乐情绪识别。
Personalized Dynamic Music Emotion Recognition with Dual-Scale Attention-Based Meta-Learning
- 采用双尺度注意力机制捕捉音乐序列的长短时依赖
- 仅需一个个性化标注样本即可预测个体情绪感知
- 在传统与个性化任务上均达到领先性能
动态音乐情绪识别(DMER)旨在预测音乐中不同片段的情绪,对音乐信息检索至关重要。现有方法难以捕捉序列数据中的长期依赖关系,且忽视个体差异对情绪感知的影响。为此,我们提出个性化动态音乐情绪识别(PDMER)问题,要求模型预测符合个人感知的情绪。为此,我们设计了基于双尺度注意力的元学习方法(DSAML)。该方法通过双尺度特征提取器融合特征,并利用双尺度注意力变换器捕捉短时与长时依赖,提升传统DMER性能。为实现PDMER,我们提出按标注者划分任务的新策略:同一任务内的样本由同一标注者标注,确保感知一致性。结合元学习,DSAML仅需一个个性化标注样本即可预测个体情绪感知。客观与主观实验表明,该方法在传统和个性化任务上均达到当前最优表现。
原文摘要 · Abstract (English)
Dynamic Music Emotion Recognition (DMER) aims to predict the emotion of different moments in music, playing a crucial role in music information retrieval. The existing DMER methods struggle to capture long-term dependencies when dealing with sequence data, which limits their performance. Furthermore, these methods often overlook the influence of individual differences on emotion perception, even though everyone has their own personalized emotional perception in the real world. Motivated by these issues, we explore more effective sequence processing methods and introduce the Personalized DMER (PDMER) problem, which requires models to predict emotions that align with personalized perception. Specifically, we propose a Dual-Scale Attention-Based Meta-Learning (DSAML) method. This method fuses features from a dual-scale feature extractor and captures both short and long-term dependencies using a dual-scale attention transformer, improving the performance in traditional DMER. To achieve PDMER, we design a novel task construction strategy that divides tasks by annotators. Samples in a task are annotated by the same annotator, ensuring consistent perception. Leveraging this strategy alongside meta-learning, DSAML can predict personalized perception of emotions with just one personalized annotation sample. Our objective and subjective experiments demonstrate that our method can achieve state-of-the-art performance in both traditional DMER and PDMER.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。