用音符符号追踪提升钢琴演奏定位精度与鲁棒性
Pairing Real-Time Piano Transcription with Symbol-level Tracking for Precise and Robust Score Following
- 将音频转为音符序列,再在符号层面匹配乐谱
- 相比纯音频方法,定位误差更小,失败率更低
- 适合需要高精度的实时音乐分析场景
实时音乐跟踪系统可实时定位演奏进度与对应乐谱位置。现有方法多依赖音频域处理,通常采用在线时间对齐技术(OLTW)结合音频表示的乐谱。尽管特征和模型策略持续优化,过去十年性能已趋于饱和。本文提出将演奏转换至符号域——即把音乐跟踪转化为符号匹配任务,即使转换存在不完美。所提系统由两个实时模块构成:音频到音符的转录模块,以及新型符号级跟踪器,连接转录结果与乐谱。与纯音频方法对比,本方法在精度(绝对误差更小)和鲁棒性(成功跟踪率更高)上均表现更优。
原文摘要 · Abstract (English)
Real-time music tracking systems follow a musical performance and at any time report the current position in a corresponding score. Most existing methods approach this problem exclusively in the audio domain, typically using online time warping (OLTW) techniques on incoming audio and an audio representation of the score. Audio OLTW techniques have seen incremental improvements both in features and model heuristics which reached a performance plateau in the past ten years. We argue that converting and representing the performance in the symbolic domain -- thereby transforming music tracking into a symbolic task -- can be a more effective approach, even when the domain transformation is imperfect. Our music tracking system combines two real-time components: one handling audio-to-note transcription and the other a novel symbol-level tracker between transcribed input and score. We compare the performance of this mixed audio-symbolic approach with its equivalent audio-only counterpart, and demonstrate that our method outperforms the latter in terms of both precision, i.e., absolute tracking error, and robustness, i.e., tracking success.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。