直接对齐音频特征与乐谱位置,精度超越传统合成音轨方法。
Precise and Simple Audio-to-Score Alignment

- 用动态规划直接匹配音频特征与乐谱符号位置
- 在钢琴独奏数据集上精度优于现有音频-音频对齐方法
- 无需单独转录模型,适配不同音色且计算高效
音频-乐谱对齐是音乐信息检索中的长期挑战,也是应用最广泛的音乐对齐任务。传统方法需将乐谱合成音频或从音频提取类似音频的特征(如钢琴卷轴)。本文提出一种直接桥接音频特征与符号级特征的新算法:通过专门设计的动态规划匹配方法,将包含起始点和频谱激活的序列音频特征与乐谱位置对齐。该方法在精度上超越广泛使用的基于合成音轨的音频-音频对齐方法,同时保持数字信号处理组件的灵活性,无需额外转录模型即可适应不同音色。其算法复杂度在符号乐谱(通常较短)和音频特征序列(通常较长)长度上最多为线性关系。文章详细描述算法并基于大规模钢琴独奏录音数据集评估对齐质量。
原文摘要 · Abstract (English)
Audio-to-score alignment is a long-standing challenge in music information retrieval and arguably the most widely applicable alignment task for music research. Alignment algorithms match two versions of a piece of music, and for this to work these versions need to be in comparable formats. Audio-to-audio alignment matches audio features; when matching audio files to scores, they must either synthesize the score or derive audio-like features by means of piano rolls or similar feature sequences. Symbolic alignment, by contrast, matches symbolically encoded notes; in an audio-to-score scenario these would be obtained by a transcription of the audio file. In this article, we present an algorithm that bridges audio-like and symbol-level features directly. Sequential audio features encoding onset and spectral activation are matched to score positions by a bespoke dynamic programming-based matching algorithm derived from symbolic alignment methods. The resulting method is both precise - surpassing widely used audio-to-audio approaches based on synthesized scores -, and remains flexible in its digital signal processing components, i.e., the method is adaptable to diverse timbral characteristics without requiring a separate transcription model. Furthermore it inherits some of the symbolic alignment runtime advantages with an algorithmic complexity that is at worst linear in the length of the (typically short) symbolic score and (typically long) audio feature sequence. In the following sections, we provide a detailed algorithm description and evaluate its alignment quality on a large-scale dataset of solo piano recordings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。