arXiv:2605.06685cs.SDeess.AS2026-05被引 1

通过音频分析生成作曲家信息论画像,揭示风格差异与创作规律。

An audio-to-analysis pipeline with certified transcription for information-theoretic profiling of the piano repertoire

论文配图:An audio-to-analysis pipeline with certified transcription for information-theoretic profiling of the piano repertoire
图 1 · 摘自论文原文
  • 基于高精度转录的音频分析流水线,从演奏录音中提取作曲家风格特征。
  • 发现古典作曲家和新古典音乐家在和声可预测性上存在显著差异。
  • 适合音乐信息检索、风格分析与作曲家研究者参考。

我们提出一个从音频到分析的流水线,生成作曲家级别的信息论风格画像,反映从聚合演奏中浮现的创作词汇。该方法基于一个在标准基准上经过验证的转录层(在MAESTRO v3.0.0测试集上F1=0.9791)。对1,238首作品及15位至少有十首作品的作曲家(涵盖巴洛克至二十世纪初)进行分析,统计和声音级的经验分布,并通过香农熵、非对称KL散度和齐普夫秩频模型进行建模。结果表明:(i) 作曲家在和声可预测性轴上有序排列,熵值范围窄(3.33–3.86比特),显示调性词汇的边际相似性;(ii) 通过最小的KL散度恢复了已知风格谱系(海顿-贝多芬、李斯特-拉赫玛尼诺夫、舒伯特-舒曼),门德尔松在此语料库中作为稳定异常点出现;(iii) 新古典主义艺术家(里希特、弗拉姆、格拉斯、阿纳尔德斯、约翰松)在转换分布的齐普夫拟合质量上明显优于历史作曲家,平均R²为0.78(新古典)对比0.46(历史)(每组≥10首作品)。该差距大于组内差异,符合极简主义倾向:使用更紧凑的转换词汇并呈现更强的频率秩规律性。所有估计均附带拉普拉斯平滑的自助法95%置信区间。

原文摘要 · Abstract (English)

We present an audio-to-analysis pipeline that produces composer-level information-theoretic profiles : reflecting compositional vocabulary as it emerges from aggregated performances : from raw recordings, built on a transcription layer whose accuracy we certify on a standard benchmark (F1 = 0.9791 on the MAESTRO v3.0.0 test set). Applied to 1,238 pieces and 15 MAESTRO composers with at least ten attributed pieces, spanning the Baroque through the early twentieth century, the pipeline derives empirical distributions over harmonic scale degrees and analyzes them through Shannon entropy, asymmetric Kullback-Leibler divergence, and Zipfian rank-frequency modeling. The resulting profiles (i) order composers along an interpretable axis of harmonic predictability, with a narrow entropy range (3.33-3.86 bits) that reveals the marginal-level similarity of tonal vocabularies; (ii) recover known stylistic lineages (Haydn-Beethoven, Liszt-Rachmaninoff, Schubert-Schumann) through the smallest KL divergences in the corpus, with Mendelssohn emerging as a stable outlier within this corpus; and (iii) separate contemporary neoclassical artists (Richter, Frahm, Glass, Arnalds, Jóhannsson) from historical composers on the quality of Zipfian fit to the transition distribution, with mean $R^2 = 0.78$ for neoclassical versus 0.46 for historical (N $\geq$ 10 pieces each). This gap is larger than the spread within either group and is consistent with a minimalist compositional tendency: a compact transition vocabulary used with sharper frequency-rank regularity than historical composers. All estimates are reported with Laplace-smoothed bootstrap 95% confidence intervals.

音乐分析信息论风格建模音频处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。