arXiv:2607.27909cs.SDcs.LG2026-07中稿 · ISMIR 2026

用音乐上下文嵌入提升钢琴演奏评估的准确性。

Integrating Contextual Embeddings into Evaluation of Expressive MIDI Piano Performances

  • 引入自监督音乐模型的上下文嵌入作为感知代理。
  • 新方法在无对齐情况下仍能捕捉演奏差异,效果优于传统指标。
  • 适合音乐生成与评估研究者使用,代码开源可复现。

现有表达性MIDI钢琴演奏的客观评估多依赖单个音符的时间、力度和时长等属性统计,但忽视了音符间的依赖关系,限制了性能相似性的判断。在生成应用中,表达属性多样性使聚合为单一指标困难。本文重新审视属性范围度量,并探索自监督符号音乐模型Aria与CLaMP3的上下文嵌入的感知特性。听觉实验表明,这些嵌入可作为感知代理,在样本级人类评分上表现与传统指标相当。为衡量条件分布相似性,我们将核音频距离(Kernel Audio Distance)适配至符号音乐领域。相比皮尔逊相关与重构误差,基于核的方法无需音符对齐,且对上下文扰动敏感。为促进可复现性,我们发布Pereval——一个开源库,集成演奏评估工具,包含属性范围与深度特征度量。

原文摘要 · Abstract (English)

Objective evaluation of expressive MIDI piano performances typically relies on attribute statistics such as timing, velocity, and duration of individual notes. However, these methods often disregard dependencies between notes, which poses a potential limitation in assessing the similarity between two sets of performances. In generative applications, the wide variety of expressive attributes makes it difficult to aggregate them into a single scalar metric for model selection. In this work, we reexamine attribute-scoped metrics and explore the perceptual properties of contextual embeddings from self-supervised symbolic music models, Aria and CLaMP3. Results from our listening study indicate that these models can be used as perceptual proxies, showing agreement with per-sample human ratings on par with traditional metrics. To measure conditional distributional similarity, we adapt Kernel Audio Distance to the symbolic music domain. Unlike Pearson correlation and reconstruction error, kernel-based methods on contextual embeddings do not require note alignment and are sensitive to contextual perturbations. To facilitate reproducibility, we release Pereval, an open-source library that integrates performance evaluation utilities, including both attribute-scoped and deep feature metrics.

音乐生成评估方法上下文嵌入MIDI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。