综述多模态音乐情感识别的最新进展与挑战
A Survey on Multimodal Music Emotion Recognition
- 提出四阶段框架:数据选择、特征提取、处理和情感预测
- 深度学习与特征融合显著提升识别效果
- 适合关注音乐推荐与治疗应用的研究者
多模态音乐情感识别(MMER)是音乐信息检索领域新兴方向,近年备受关注。本文系统梳理了该领域的最新进展,提出包含多模态数据选择、特征提取、特征处理和情感预测四个阶段的框架。分析表明,深度学习方法取得显著进步,特征融合技术日益重要。然而,仍面临大规模标注数据稀缺、多模态数据集不足及实时处理能力欠缺等挑战。论文还指出了当前研究的关键空白,并提出未来发展方向,强调开发鲁棒、可扩展、可解释模型的重要性,对音乐推荐系统、治疗工具和娱乐应用具有重要意义。
原文摘要 · Abstract (English)
Multimodal music emotion recognition (MMER) is an emerging discipline in music information retrieval that has experienced a surge in interest in recent years. This survey provides a comprehensive overview of the current state-of-the-art in MMER. Discussing the different approaches and techniques used in this field, the paper introduces a four-stage MMER framework, including multimodal data selection, feature extraction, feature processing, and final emotion prediction. The survey further reveals significant advancements in deep learning methods and the increasing importance of feature fusion techniques. Despite these advancements, challenges such as the need for large annotated datasets, datasets with more modalities, and real-time processing capabilities remain. This paper also contributes to the field by identifying critical gaps in current research and suggesting potential directions for future research. The gaps underscore the importance of developing robust, scalable, a interpretable models for MMER, with implications for applications in music recommendation systems, therapeutic tools, and entertainment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。