arXiv:2507.04230cs.SDcs.AI2025-07被引 2

从钢琴音频中连续估算踏板深度,提升音乐表现力建模精度

High-Resolution Sustain Pedal Depth Estimation from Piano Audio Across Room Acoustics

  • 采用Transformer架构,实现踏板深度的连续值预测
  • 在多种混响环境下测试,模型仍保持较高估计准确率
  • 适合音乐信息检索与自动伴奏系统开发者参考

以往钢琴延音踏板检测多为二分类任务,限制了其在真实演奏场景中的应用,因踏板深度显著影响音乐表现力。本文提出一种高分辨率连续踏板深度估计方法,采用基于Transformer的架构,在传统二分类任务上达到顶尖水平,并实现精确的连续值预测。该方法能捕捉细腻的音乐表达,而基线模型因仅输出开关状态难以胜任。研究还通过包含不同声学条件的合成数据集,探究混响对踏板估计的影响。训练时使用多种房间设置,测试时采用“留一法”评估未见环境下的表现。结果表明,三种模型均对未知声学环境不鲁棒;统计分析显示混响显著影响预测结果,导致系统性高估。

原文摘要 · Abstract (English)

Piano sustain pedal detection has previously been approached as a binary on/off classification task, limiting its application in real-world piano performance scenarios where pedal depth significantly influences musical expression. This paper presents a novel approach for high-resolution estimation that predicts continuous pedal depth values. We introduce a Transformer-based architecture that not only matches state-of-the-art performance on the traditional binary classification task but also achieves high accuracy in continuous pedal depth estimation. Furthermore, by estimating continuous values, our model provides musically meaningful predictions for sustain pedal usage, whereas baseline models struggle to capture such nuanced expressions with their binary detection approach. Additionally, this paper investigates the influence of room acoustics on sustain pedal estimation using a synthetic dataset that includes varied acoustic conditions. We train our model with different combinations of room settings and test it in an unseen new environment using a "leave-one-out" approach. Our findings show that the two baseline models and ours are not robust to unseen room conditions. Statistical analysis further confirms that reverberation influences model predictions and introduces an overestimation bias.

音频分析持续估计音乐表现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。