只需点击重复记号,就能高效对齐真实演出音频与乐谱图像。
Just Label the Repeats for In-The-Wild Audio-to-Score Alignment
- 用户仅需点击重复记号标注跳转,大幅减少人工标注量。
- 结合改进特征表示,对齐准确率提升150%(33%→82%)。
- 适合音乐信息检索、自动伴奏生成等场景的开发者使用。
本文提出一种高效的离线音频-乐谱对齐流程,用于真实演出音频与乐谱扫描图像的高质量匹配。现有基于动态时间规整(DTW)的方法虽无需人工标注,但对重复记号引起的跳转处理效果不佳。为此,我们设计了一种工作流和交互界面,仅需用户点击重复记号即可快速标注跳转位置,少量人工干预下显著提升对齐质量。同时,通过两项改进进一步优化:(1) 在乐谱特征中融合小节检测信息;(2) 使用音乐转录模型输出的原始击键概率,替代传统钢琴卷帘表示。我们还提出一种以小节为单位衡量对齐误差的评估协议。实验表明,该方法在新评估标准下相较先前工作,对齐准确率相对提升150%(从33%提高至82%)。
原文摘要 · Abstract (English)
We propose an efficient workflow for high-quality offline alignment of in-the-wild performance audio and corresponding sheet music scans (images). Recent work on audio-to-score alignment extends dynamic time warping (DTW) to be theoretically able to handle jumps in sheet music induced by repeat signs-this method requires no human annotations, but we show that it often yields low-quality alignments. As an alternative, we propose a workflow and interface that allows users to quickly annotate jumps (by clicking on repeat signs), requiring a small amount of human supervision but yielding much higher quality alignments on average. Additionally, we refine audio and score feature representations to improve alignment quality by: (1) integrating measure detection into the score feature representation, and (2) using raw onset prediction probabilities from a music transcription model instead of piano roll. We propose an evaluation protocol for audio-to-score alignment that computes the distance between the estimated and ground truth alignment in units of measures. Under this evaluation, we find that our proposed jump annotation workflow and improved feature representations together improve alignment accuracy by 150% relative to prior work (33% to 82%).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。