用事件编码音乐,兼顾效率与结构信息。
Pianoroll-Event: A Novel Score Representation for Symbolic Music
- 将乐谱转为四种事件:帧、间隙、模式、结构,融合时空特性。
- 编码效率提升1.36到7.16倍,序列更短词汇更少。
- 适合音乐生成、建模研究者,尤其关注效率与结构的场景。
符号化音乐表示是计算音乐学中的基础挑战。基于网格的表示虽能保持音高-时间空间对应关系,但存在数据稀疏问题,导致编码效率低;离散事件表示虽编码紧凑,却难以捕捉结构不变性和空间局部性。为此,我们提出Pianoroll-Event,一种新型编码方案,通过事件描述钢琴卷谱表示,兼顾结构属性与编码效率,同时保留时间依赖与局部空间模式。具体设计四种互补事件类型:帧事件(用于时间边界)、间隙事件(用于稀疏区域)、模式事件(用于音符模式)和音乐结构事件(用于音乐元数据)。该方法在序列长度与词表大小间取得有效平衡,在多种自回归架构上的实验表明,使用该表示的模型在定量与人工评估中均持续优于基线。
原文摘要 · Abstract (English)
Symbolic music representation is a fundamental challenge in computational musicology. While grid-based representations effectively preserve pitch-time spatial correspondence, their inherent data sparsity leads to low encoding efficiency. Discrete-event representations achieve compact encoding but fail to adequately capture structural invariance and spatial locality. To address these complementary limitations, we propose Pianoroll-Event, a novel encoding scheme that describes pianoroll representations through events, combining structural properties with encoding efficiency while maintaining temporal dependencies and local spatial patterns. Specifically, we design four complementary event types: Frame Events for temporal boundaries, Gap Events for sparse regions, Pattern Events for note patterns, and Musical Structure Events for musical metadata. Pianoroll-Event strikes an effective balance between sequence length and vocabulary size, improving encoding efficiency by 1.36\times to 7.16\times over representative discrete sequence methods. Experiments across multiple autoregressive architectures show models using our representation consistently outperform baselines in both quantitative and human evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。