用声音参数反推不同乐器的力度值,让音乐生成更自然
Beyond Piano: Cross-Instrument MIDI Velocity Estimation via Differentiable SoundFont Proxies

- 用可微分音色库代理监督力度预测,聚焦于响度相关声学特征
- 在钢琴和吉他上实现跨乐器力度估计,效果优于波形重建方法
- 适合做音乐生成、演奏分析的研究者和开发者
许多音乐数据集包含MIDI音符但缺乏可靠的力度值,通常默认为常数。这一缺失在钢琴以外的乐器领域尤为严重,因为力度是表达性演奏、音乐生成和性能分析的核心。本文研究在标签稀缺情况下跨乐器的MIDI力度估计。从训练好的钢琴力度估计器出发,将目标乐器适配问题转化为预测与音源渲染匹配的力度值。该过程可通过可微分合成器(Diff-Synth)或本文提出的可微分SoundFont代理(Diff-SFProxy)驱动。Diff-SFProxy通过音符级响度相关声学参数监督力度,而非波形重建,使梯度聚焦于力度相关行为。在钢琴和吉他上的实验表明,Diff-SFProxy在跨乐器力度估计中表现有效,而基于波形域的Diff-Synth则导致性能下降。
原文摘要 · Abstract (English)
Many music datasets contain MIDI notes but lack reliable velocities, defaulting to a constant value. This absence is especially problematic outside the piano domain, as velocity is a core component for expressive rendering, music generation, and performance analysis. This paper studies cross-instrument MIDI velocity estimation in this label-scarce setting. Starting from a piano-trained velocity estimator, we recast target-instrument adaptation as predicting renderer-conditioned velocities whose rendering matches the dynamics of the performance audio. This adaptation can be driven by either differentiable synthesizers (Diff-Synth) or our proposed differentiable SoundFont proxies (Diff-SFProxy). We highlight the Diff-SFProxy: it supervises velocity through note-wise, loudness-related acoustic parameters rather than waveform reconstruction, focusing gradients on velocity-dependent behavior. Experiments on piano and guitar show that Diff-SFProxy is effective for cross-instrument MIDI velocity estimation, while waveform-domain Diff-Synth degrades performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。