贝多芬月光奏鸣曲结构暗合机器学习机制,无需类比
Moonlight in Latent Space: Chirality and Structural Correspondence Between Beethoven's Op. 27 No. 2 and Machine Learning Mechanisms

- 用音符结构对应三种神经网络架构,非比喻而是数学映射
- 无乐理输入下聚类自动还原调性结构,且音乐有强顺序依赖
- 音乐顺序信息高于噪声,但语序约束比音乐更强(高手性)
我们发现贝多芬《月光奏鸣曲》(作品27之2)的三个乐章在结构上分别对应三种不同的机器学习架构——不是通过类比,而是通过结构性对应。通过对乐谱进行计算分析(熵、Jensen-Shannon散度、不协和度、手部分布重叠、自相似矩阵、时间记忆衰减、上下文音高嵌入),我们得出四个反直觉发现:(1)听感“温度”由吞吐量决定,而非分布宽度;(2)最轻盈的乐章具有最高不协和度;(3)三乐章分别实现流式、循环与周期性位置编码记忆架构;(4)同一音级在不同乐章中具有不同上下文身份,类似NLP中的上下文嵌入与静态嵌入,并且无音乐理论输入的无监督聚类可恢复调性结构。我们构建了逆向声学化(将分析特征解码回MIDI),并量化了编码-解码循环的左右手性:哪些分布保持不变,而序列顺序则被破坏。受听众观察‘听起来像无法叠加的镜像异构体’启发,手性测量显示重构损失随n元语法阶数单调上升。自助法基线与子采样检验确认所有乐章的序列信息均高于噪声水平,尽管原始数值受样本量干扰。跨域比较显示自然语言的手性高于音乐,反映更强的序列约束。
原文摘要 · Abstract (English)
We show that the three movements of Beethoven's "Moonlight Sonata" (Op. 27 No. 2) instantiate three distinct machine learning architectures -- not by analogy, but by structural correspondence. Through computational analysis of the score (entropy, Jensen-Shannon divergence, dissonance, hand distributional overlap, self-similarity matrices, temporal memory decay, and contextual pitch embeddings), we establish four counterintuitive findings: (1) perceived musical "temperature" is governed by throughput, not distributional width; (2) the lightest movement carries the highest dissonance; (3) the movements implement streaming, recurrent, and periodic positional encoding memory architectures; and (4) the same pitch class acquires different contextual identities across movements, analogous to contextual vs.static embeddings in NLP -- and unsupervised clustering recovers the tonal structure without music-theoretic input. We construct a reverse sonification (decoding analytical features back into MIDI) and quantify the chirality of the encode-decode cycle: what distributions preserve and sequential ordering destroys. Prompted by a listener's observation that the decoded piece sounds like "mirror isomers that can't be superimposed," the chirality measurement reveals reconstruction loss increasing monotonically with n-gram order. Bootstrap baselines and subsample checks confirm all movements carry sequential information above noise, though raw values are confounded by sample size. Cross-domain comparison shows natural language has higher chirality than music, reflecting stronger sequential constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。