用小模型同时预测钢琴力度与节拍结构,效率更高。
Joint Estimation of Piano Dynamics and Metrical Structure with a Multi-task Multi-Scale Network
- 多任务多尺度网络共享特征表示,联合预测力度、变化点、节拍和强拍。
- 输入采用Bark尺度响度特征,模型大小仅0.5M,支持60秒长音频处理。
- 在MazurkaBL数据集上四项指标均达当前最优,适合大规模音乐表达分析。
从音频中估计钢琴力度是计算音乐分析中的基础挑战。本文提出一种高效的多任务网络,通过共享潜在表示联合预测力度水平、力度变化点、节拍和强拍,这些目标共同构成乐谱中力度的节律结构。受近期声乐力度研究启发,采用多尺度网络作为主干,输入为Bark尺度特定响度特征。相比传统log-Mel输入,模型规模从14.7M降至0.5M,支持长序列输入。音频分段长度达60秒,是常规节拍追踪长度的两倍。在公开的MazurkaBL数据集上,该模型在所有任务上均取得当前最佳性能。本工作为钢琴力度估计设立了新基准,提供了一个强大且紧凑的工具,推动了大规模、资源高效音乐表现分析的发展。
原文摘要 · Abstract (English)
Estimating piano dynamic from audio recordings is a fundamental challenge in computational music analysis. In this paper, we propose an efficient multi-task network that jointly predicts dynamic levels, change points, beats, and downbeats from a shared latent representation. These four targets form the metrical structure of dynamics in the music score. Inspired by recent vocal dynamic research, we use a multi-scale network as the backbone, which takes Bark-scale specific loudness as the input feature. Compared to log-Mel as input, this reduces model size from 14.7 M to 0.5 M, enabling long sequential input. We use a 60-second audio length in audio segmentation, which doubled the length of beat tracking commonly used. Evaluated on the public MazurkaBL dataset, our model achieves state-of-the-art results across all tasks. This work sets a new benchmark for piano dynamic estimation and delivers a powerful and compact tool, paving the way for large-scale, resource-efficient analysis of musical expression.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。