arXiv:2601.19029cs.SDeess.AS2026-01

用音频模型评估钢琴演奏,比传统符号方法更准

Audio Foundation Models Outperform Symbolic Representations for Piano Performance Evaluation

  • 用预训练音频模型直接分析演奏音色与表现力
  • 在19个感知维度上,音频模型效果提升55%
  • 适合音乐人工智能、自动评分系统研究者

自动钢琴演奏评价传统依赖符号(MIDI)表示,仅捕捉音符信息,忽略体现表现力的声学细节。本文提出使用预训练音频基础模型(MuQ和MERT)预测19个感知维度的演奏质量。基于PercePiano MIDI文件合成的音频(使用Pianoteq音源),在相同数据源下对比音频与符号方法。最佳模型(MuQ层9-12,配合Pianoteq音源增强)达到R² = 0.537(95%置信区间:[0.465, 0.575]),相比符号基线(R² = 0.347)提升55%,统计显著(p < 10⁻²⁵),且在全部19个维度上均更优。通过跨音源泛化验证(R² = 0.534 ± 0.075)、外部数据集难度相关性分析(rho = 0.623)及多演奏者一致性检验,证实方法有效性。音频-符号融合分析显示误差高度相关(r = 0.738),说明融合无实质增益:音频表示已足够。论文发布完整训练流程、预训练模型与推理代码。

原文摘要 · Abstract (English)

Automated piano performance evaluation traditionally relies on symbolic (MIDI) representations, which capture note-level information but miss the acoustic nuances that characterize expressive playing. I propose using pre-trained audio foundation models, specifically MuQ and MERT, to predict 19 perceptual dimensions of piano performance quality. Using synthesized audio from PercePiano MIDI files (rendered via Pianoteq), I compare audio and symbolic approaches under controlled conditions where both derive from identical source data. The best model, MuQ layers 9-12 with Pianoteq soundfont augmentation, achieves R^2 = 0.537 (95% CI: [0.465, 0.575]), representing a 55% improvement over the symbolic baseline (R^2 = 0.347). Statistical analysis confirms significance (p < 10^-25) with audio outperforming symbolic on all 19 dimensions. I validate the approach through cross-soundfont generalization (R^2 = 0.534 +/- 0.075), difficulty correlation with an external dataset (rho = 0.623), and multi-performer consistency analysis. Analysis of audio-symbolic fusion reveals high error correlation (r = 0.738), explaining why fusion provides minimal benefit: audio representations alone are sufficient. I release the complete training pipeline, pretrained models, and inference code.

音频模型音乐评估表现力分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。