arXiv:2504.05009cs.SDcs.IR2025-04被引 1

用机器学习拆解20位爵士钢琴家的风格,准确率达94%

Deconstructing Jazz Piano Style Using Machine Learning

  • 构建多输入模型,分别分析旋律、和声、节奏与力度四个音乐维度
  • 在84小时录音上实现20位钢琴家识别94%准确率,超越现有水平
  • 开源模型与网页工具,适合音乐研究者与风格分析爱好者

艺术风格研究已有数百年历史,近年来机器学习为计算理解风格提供了新可能。然而,如何确保模型输出契合实践者与评论家的关注点仍是重大挑战。本文聚焦音乐风格,借助其深厚的理论与数学分析传统,训练多种监督学习模型,在精心整理的84小时录音数据集上识别20位标志性爵士钢琴家,并解析其决策过程。模型采用新型多输入架构,可独立分析旋律、和声、节奏与力度四个音乐维度。该方法不仅回答了音乐理论中的基础问题,还实现了音乐表演者识别的最新性能(20类中达94%准确率)。我们开源了模型代码及配套网页应用,供探索音乐风格。

原文摘要 · Abstract (English)

Artistic style has been studied for centuries, and recent advances in machine learning create new possibilities for understanding it computationally. However, ensuring that machine-learning models produce insights aligned with the interests of practitioners and critics remains a significant challenge. Here, we focus on musical style, which benefits from a rich theoretical and mathematical analysis tradition. We train a variety of supervised-learning models to identify 20 iconic jazz musicians across a carefully curated dataset of 84 hours of recordings, and interpret their decision-making processes. Our models include a novel multi-input architecture that enables four musical domains (melody, harmony, rhythm, and dynamics) to be analysed separately. These models enable us to address fundamental questions in music theory and also advance the state-of-the-art in music performer identification (94% accuracy across 20 classes). We release open-source implementations of our models and an accompanying web application for exploring musical styles.

风格分析音乐生成机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。