改进钢琴转录模型,更准且更小,适合音乐信号精细分析。
Improved Architecture for High-resolution Piano Transcription to Efficiently Capture Acoustic Characteristics of Music Signals

- 用恒Q变换提取音频特征,更好捕捉音乐时频特性。
- 在注释精度上全面超越基线,模型体积显著减小。
- 适合需要高效高精度转录的音乐信息处理场景。
自动音乐转录(AMT)旨在将音乐信号转换为乐谱,是音乐信息检索的重要任务。近期研究采用高分辨率标签(如音符连续起始与终止时间)作为训练目标,显著提升了转录性能。然而仍存在谐波被误识别为假阳性音符、模型规模过大等问题。为此,本文提出一种改进的高分辨率钢琴转录模型,以更好地捕捉音乐信号的声学特征。首先,采用恒Q变换(Constant-Q Transform)作为输入表示,更适配音乐信号;其次设计两种架构:一是基于空洞卷积的卷积循环神经网络(CRNN),二是结合CRNN与非自回归Transformer解码器的编码器-解码器结构。系统实验表明,相较于基线高分辨率AMT系统,所提模型在音符级指标上实现持续提升,且模型规模显著缩小,为未来工作提供重要参考。
原文摘要 · Abstract (English)
Automatic music transcription (AMT), aiming to convert musical signals into musical notation, is one of the important tasks in music information retrieval. Recently, previous works have applied high-resolution labels, i.e., the continuous onset and offset times of piano notes, as training targets, achieving substantial improvements in transcription performance. However, there still remain some issues to be addressed, e.g., the harmonics of notes are sometimes recognized as false positive notes, and the size of AMT model tends to be larger to improve the transcription performance. To address these issues, we propose an improved high-resolution piano transcription model to well capture specific acoustic characteristics of music signals. First, we employ the Constant-Q Transform as the input representation to better adapt to musical signals. Moreover, we have designed two architectures: the first is based on a convolutional recurrent neural network (CRNN) with dilated convolution, and the second is an encoder-decoder architecture that combines CRNN with a non-autoregressive Transformer decoder. We conduct systematic experiments for our models. Compared to the high-resolution AMT system used as a baseline, our models effectively achieve 1) consistent improvement in note-level metrics, and 2) the significant smaller model size, which shed lights on future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。