Harmonica用多深度谐波卷积实现高精度轻量级音乐转录,适合实时应用。
Harmonica: Accurate and Lightweight Instrument-Agnostic Music Transcription

- 基于多深度谐波卷积构建,不依赖具体乐器
- 纳米版仅26.3K参数,推理速度达1622.5倍实时,帧F1达0.796
- 在多个数据集上优于基础模型,适合资源受限场景
本文提出Harmonica,一套基于多深度谐波卷积的无乐器依赖音乐转录模型。在各规模模型中,Harmonica均表现最优:x-large模型达到当前最佳性能;medium版本在准确率与推理速度上均优于所有基线;nano版本仅含26.3K参数,推理速度达1,622.5倍实时,在开发集上帧F1达0.796,较Basic Pitch提升14.6个百分点。通过对比实验验证,多深度谐波卷积能有效利用谐波信息,显著提升转录性能,优于谐波堆叠、谐波注意力、单深度谐波卷积及HD-Conv层等方法。
原文摘要 · Abstract (English)
This paper introduces Harmonica, a family of instrument-agnostic music transcription models built around multi-depth harmonic convolution. At each model scale, Harmonica achieves the best performance among the evaluated models: the x-large model attains state-of-the-art performance in instrument-agnostic transcription, while the medium variant offers competitive accuracy with faster inference than all baselines. Pushing the limit of computational efficiency, the nano variant has only 26.3K parameters and runs at 1,622.5 times real time, yet achieves a frame F1 of 0.796 on the development set, 14.6 percentage points higher than Basic Pitch. We further demonstrate that multi-depth harmonic convolution effectively exploits harmonic information to benefit transcription performance through comparative experiments with existing harmonic aggregation methods, including harmonic stacking, harmonic attention, single-depth harmonic convolution, and the HD-Conv layer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。