arXiv:2504.05364cs.SDcs.AI2025-04

提出新位置编码方法,提升音乐生成效率与质量。

Of All StrIPEs: Investigating Structure-informed Positional Encoding for Efficient Music Generation

  • 基于核方法统一分析旋转与随机傅里叶位置编码。
  • 新方法RoPEPool在旋律和声任务中超越现有模型。
  • 适合关注音乐生成与高效位置编码的研究者。

尽管音乐生成仍是生成模型(如Transformer)的挑战领域,但近期研究发现,在位置编码中引入音乐结构信息,并结合基于随机傅里叶特征(RFF)的核近似技术,可将计算成本从二次方降低至线性。然而,此类RFF-based高效位置编码与基于旋转矩阵的旋转位置编码(RoPE)之间的性能对比尚不清晰。本文提出一个基于核方法的统一框架,用于分析这两类高效位置编码。利用该框架,我们开发出一种名为RoPEPool的新位置编码方法,能够从时间序列中提取因果关系。通过符号化音乐生成任务——旋律和声——的实证验证,结果表明:结合高信息量的结构先验,RoPEPool优于所有对比方法。

原文摘要 · Abstract (English)

While music remains a challenging domain for generative models like Transformers, a two-pronged approach has recently proved successful: inserting musically-relevant structural information into the positional encoding (PE) module and using kernel approximation techniques based on Random Fourier Features (RFF) to lower the computational cost from quadratic to linear. Yet, it is not clear how such RFF-based efficient PEs compare with those based on rotation matrices, such as Rotary Positional Encoding (RoPE). In this paper, we present a unified framework based on kernel methods to analyze both families of efficient PEs. We use this framework to develop a novel PE method called RoPEPool, capable of extracting causal relationships from temporal sequences. Using RFF-based PEs and rotation-based PEs, we demonstrate how seemingly disparate PEs can be jointly studied by considering the content-context interactions they induce. For empirical validation, we use a symbolic music generation task, namely, melody harmonization. We show that RoPEPool, combined with highly-informative structural priors, outperforms all methods.

音乐生成位置编码Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。