不同音乐编码方式影响网络结构与感知鲁棒性,关键在信息密度与可预测性权衡。
Trade-offs between structural richness and perceptual robustness in music network representations
- 用节点和边构建音乐网络,比较八种编码对状态空间的重塑作用
- 单特征编码熵高但误差低,多特征编码更精细但易受感知约束干扰
- 预测性流动中局部意外集中,适合研究音乐认知与生成模型设计
音乐是时间序列中具有结构和感知丰富性的声音组合,其感知由对后续内容的预期与不确定性共同塑造。然而,我们从音乐中推断出的不确定性取决于音乐被编码为事件序列的方式。本文采用网络表示法,将事件类型作为节点,观测到的转移作为有向边,系统分析八种钢琴音乐编码(从单一特征词汇到多特征组合)如何改变所恢复的转移结构及其在模拟感知约束下的鲁棒性。这些表征选择重构了状态空间,根本性地改变了网络拓扑,调整了不确定性在转移间的分布。通过引入感知约束模型以捕捉对转移统计信息的不完全获取,结果表明:压缩的单特征表示生成密集的转移结构,具有更高的熵率(平均每步不确定性更高),但模型误差较低,说明受限估计仍贴近真实语料转移;而更丰富的多特征表示虽保留更细粒度差异,却扩大状态空间,锐化转移特征、降低熵率,并增加模型误差。整体上,不确定性集中在扩散中心节点,而模型误差在此处仍较低,表明存在可预测流动与局部突变共存的信息景观。结果表明,特征选择不仅影响重建网络,还决定了可用的转移统计量及其在感知约束下的脆弱性。
原文摘要 · Abstract (English)
Music is a structured and perceptually rich sequence of sounds in time, whose perception is shaped by the interplay of expectation and uncertainty about what comes next. Yet the uncertainty we infer from music depends on how the musical piece is encoded as an event sequence. In this work, we use network representations, in which event types are nodes and observed transitions are directed edges, to compare how different feature encodings shape the transition structure we recover and how robust that structure is under modeled perceptual constraints. We systematically analyse eight encodings of piano music, from single-feature vocabularies to richer multi-feature combinations. These representational choices reorganize the state space and fundamentally reshape network topology, shifting how uncertainty is distributed across transitions. To connect these descriptive differences to perception, we adopt a perceptual-constraint model that captures imperfect access to transition statistics. Overall, compressed single-feature representations yield dense transition structures with higher entropy rates, corresponding to higher average uncertainty per step, yet low model error, indicating that the constrained estimate stays close to the corpus transitions. In contrast, richer multi-feature representations preserve finer distinctions but expand the state space, sharpen transition profiles, lower entropy rates, and increase model error. Finally, across representations, uncertainty concentrates in diffusion-central nodes while model error remains low there, suggesting an informational landscape in which predictable flow coexists with localized surprise. Overall, our results show that feature choice shapes not only the networks we reconstruct, but also which transition statistics are available and how vulnerable they are to distortion under perceptual constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。