用监督学习提升游戏音乐结构分割准确率
Supervised Learning for Game Music Segmentation
- 结合卷积与循环神经网络进行音乐结构分割
- 在309段标注数据上达到顶尖无监督方法水平
- 为游戏音乐生成提供可解释的结构基础
当前基于神经网络的模型,包括Transformer,在处理统一重复的音乐素材时,难以生成令人难忘且易于理解的音乐,主要因为缺乏对音乐结构的理解。因此,这些模型在游戏行业中很少被采用。许多学者认为,建模音乐结构可使模型在更高层次上获得理解,从而提升音乐生成质量。本研究旨在探索监督学习方法在结构分割任务中的表现,这是音乐结构建模的第一步。研究构建了一个包含309个结构标注的游戏音乐音频数据集,提出的方法结合了卷积神经网络与循环神经网络,在更少训练资源下达到了与当前最优无监督方法相当的性能。
原文摘要 · Abstract (English)
At present, neural network-based models, including transformers, struggle to generate memorable and readily comprehensible music from unified and repetitive musical material due to a lack of understanding of musical structure. Consequently, these models are rarely employed by the games industry. It is hypothesised by many scholars that the modelling of musical structure may inform models at a higher level, thereby enhancing the quality of music generation. The aim of this study is to explore the performance of supervised learning methods in the task of structural segmentation, which is the initial step in music structure modelling. An audio game music dataset with 309 structural annotations was created to train the proposed method, which combines convolutional neural networks and recurrent neural networks, achieving performance comparable to the state-of-the-art unsupervised learning methods with fewer training resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。