arXiv:2501.07474cs.SDcs.AI2025-01中稿 · the 2025 IEEE Inte…被引 3

用音频自回归模型预测音乐意外性,关联人脑反应与音色变化。

Estimating Musical Surprisal in Audio

  • 用Transformer预测压缩音频特征,以信息量衡量音乐意外性。
  • 后期段落信息量更高,与听觉预期偏差一致。
  • 能预测脑电反应,适合音乐认知与神经科学研究者。

在计算建模音乐意外性时,已有研究采用自回归模型对符号音乐进行一步预测的信息内容(IC)作为意外性的代理指标。本研究探索该方法在音乐音频上的适用性:训练一个自回归Transformer模型,预测预训练自编码器网络的压缩音频表征。通过重复实验验证学习效果,发现信息量随重复降低。分析不同段落类型(如A段或B段)的平均信息量,发现出现在作品后期的段落平均信息量更高。进一步分析信息量与音频及音乐特征的关系,发现其与音色变化、响度显著相关,与不协和度、节奏复杂度、起始密度等有一定关联。最后,检验信息量能否预测人脑对歌曲的脑电反应,验证其建模人类音乐意外性的潜力。代码已公开于github.com/sonycslparis/audioic。

原文摘要 · Abstract (English)

In modeling musical surprisal expectancy with computational methods, it has been proposed to use the information content (IC) of one-step predictions from an autoregressive model as a proxy for surprisal in symbolic music. With an appropriately chosen model, the IC of musical events has been shown to correlate with human perception of surprise and complexity aspects, including tonal and rhythmic complexity. This work investigates whether an analogous methodology can be applied to music audio. We train an autoregressive Transformer model to predict compressed latent audio representations of a pretrained autoencoder network. We verify learning effects by estimating the decrease in IC with repetitions. We investigate the mean IC of musical segment types (e.g., A or B) and find that segment types appearing later in a piece have a higher IC than earlier ones on average. We investigate the IC's relation to audio and musical features and find it correlated with timbral variations and loudness and, to a lesser extent, dissonance, rhythmic complexity, and onset density related to audio and musical features. Finally, we investigate if the IC can predict EEG responses to songs and thus model humans' surprisal in music. We provide code for our method on github.com/sonycslparis/audioic.

音乐意外性自回归模型音频表征脑电反应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。