arXiv:2604.22290cs.SDcs.MM2026-04中稿 · the 5th Internatio…

用Transformer模型基于节拍信息精准量化MIDI演奏,效果优于现有方法。

Transformer-Based Rhythm Quantization of Performance MIDI Using Beat Annotations

  • 采用Transformer架构融合乐谱与演奏数据,利用节拍标注实现同步对齐。
  • 在ASAP数据集上达到97.3%的起音F1值和83.3%的音符时值准确率。
  • 支持跨调式泛化,适合音乐信息检索与自动记谱系统开发者使用。

节奏转录是记谱级自动音乐转录(AMT)的关键子任务。尽管深度学习模型已广泛用于音频和MIDI演奏中度量网格的检测,但基于节拍的节奏量化仍研究不足。本文提出一种利用先验节拍信息进行MIDI演奏量化的新型深度学习方法。该方法基于Transformer架构,有效处理乐谱与演奏数据的同步信息,核心包括数据准备、基于节拍的预量化对齐机制,以及针对此任务定制的MIDI分词器。我们基于T5架构改进模型以满足节奏量化需求,并采用评分级指标进行客观评估。通过系统性优化数据表示与模型结构,结合演奏端时间抖动、移调、音符删除等增强策略提升鲁棒性。定性分析显示,本模型在多个曲例上优于现有概率与深度学习模型。在ASAP数据集上,起音F1得分为97.3%,音符时值准确率达83.3%。模型在未训练过的调式间具有良好泛化能力,生成可读乐谱。在乐器特异性数据集上微调后,性能进一步提升,捕捉到典型的节奏与旋律特征。本工作为基于节拍的MIDI量化提供了一个稳健且灵活的Transformer框架。

原文摘要 · Abstract (English)

Rhythm transcription is a key subtask of notation-level Automatic Music Transcription (AMT). While deep learning models have been extensively used for detecting the metrical grid in audio and MIDI performances, beat-based rhythm quantization remains largely unexplored. In this work, we introduce a novel deep learning approach for quantizing MIDI performances using a priori beat information. Our method leverages the transformer architecture to effectively process synchronized score and performance data for training a quantization model. Key components of our approach include dataset preparation, a beat-based pre-quantization method to align performance and score times within a unified framework, and a MIDI tokenizer tailored for this task. We adapt a transformer model based on the T5 architecture to meet the specific requirements of rhythm quantization. The model is evaluated using a set of score-level metrics designed for objective assessment of quantization performance. Through systematic evaluation, we optimize both data representation and model architecture. Additionally, we apply performance and score augmentations, such as transposition, note deletion, and performance-side time jitter, to enhance the model's robustness. Finally, a qualitative analysis compares our model's quantization performance against state-of-the-art probabilistic and deep-learning models on various example pieces. Our model achieves an onset F1-score of 97.3% and a note value accuracy of 83.3% on the ASAP dataset. It generalizes well across time signatures, including those not seen during training, and produces readable score output. Fine-tuning on instrument-specific datasets further improves performance by capturing characteristic rhythmic and melodic patterns. This work contributes a robust and flexible framework for beat-based MIDI quantization using transformer models.

节奏量化TransformerMIDI处理自动记谱

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。