用变压器模型直接分析性能MIDI,精准追踪节拍与强拍。
Beat and Downbeat Tracking in Performance MIDI Using an End-to-End Transformer Architecture
- 端到端Transformer架构,将MIDI序列转为节拍标注。
- 在4个数据集上超越现有符号音乐节拍追踪方法,F1分数领先。
- 适合音乐信息检索、自动乐谱生成等场景的开发者参考。
性能MIDI中的节拍追踪是记谱级音乐转录与节奏分析的重要挑战,但现有方法多集中于音频输入。本文提出一种基于端到端Transformer架构的节拍与强拍追踪模型,采用编码器-解码器结构实现从MIDI输入到节拍标注的序列到序列转换。我们引入新颖的数据预处理技术,包括动态增强与优化的分词策略,提升模型在不同数据集上的准确率与泛化能力。在A-MAPS、ASAP、GuitarSet和Leduc数据集上进行大量实验,对比了当前最优的隐马尔可夫模型(HMM)与深度学习方法。结果表明,该模型在多种音乐风格与乐器下均优于现有符号音乐节拍追踪方法,取得具有竞争力的F1分数。研究证实了变压器架构在符号节拍追踪中的潜力,并建议未来与自动音乐转录系统结合,以增强音乐分析与乐谱生成能力。
原文摘要 · Abstract (English)
Beat tracking in musical performance MIDI is a challenging and important task for notation-level music transcription and rhythmical analysis, yet existing methods primarily focus on audio-based approaches. This paper proposes an end-to-end transformer-based model for beat and downbeat tracking in performance MIDI, leveraging an encoder-decoder architecture for sequence-to-sequence translation of MIDI input to beat annotations. Our approach introduces novel data preprocessing techniques, including dynamic augmentation and optimized tokenization strategies, to improve accuracy and generalizability across different datasets. We conduct extensive experiments using the A-MAPS, ASAP, GuitarSet, and Leduc datasets, comparing our model against state-of-the-art hidden Markov models (HMMs) and deep learning-based beat tracking methods. The results demonstrate that our model outperforms existing symbolic music beat tracking approaches, achieving competitive F1-scores across various musical styles and instruments. Our findings highlight the potential of transformer architectures for symbolic beat tracking and suggest future integration with automatic music transcription systems for enhanced music analysis and score generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。