arXiv:2607.19776cs.SDcs.AI2026-07

用心理感知的节奏音高单元生成更连贯的长段旋律

RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling

论文配图:RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling
图 1 · 摘自论文原文
  • 以可变长度的节奏音高原语为单位,模拟人类对乐句的感知分组
  • 生成旋律在长期结构和音乐性上均优于传统基于小节的方法
  • 融合音乐心理学机制,适合关注音乐感知一致性的研究者

现有符号化音乐生成模型通常以小节作为基本结构单元,但人类对乐句的感知常与记谱小节线不一致,导致长程结构碎片化。本文提出RPPNet——一种两阶段深度学习架构,采用可变结构边界。首先生成可变长度的节奏-音高原语(RPP)序列,每个RPP编码音符数量、节奏模式与音高轮廓;随后将RPP序列解码为具体音符。RPP的分组由声学线索、听觉惯性及相似性感知等音乐心理学机制自动推导。实验表明,RPPNet生成的旋律在长期结构与音乐性方面均显著优于基线模型,所有主观评价维度均有提升。消融实验确认性能提升源于心理表征的结构正确性,而非模型容量。该工作为音乐生成提供了跨学科视角,融合音乐理论、计算建模与音乐心理学。

原文摘要 · Abstract (English)

Existing symbolic music generation models typically use bars as the basic structural unit. However, human perception of musical phrases often does not align with notated bar lines, leading to long-term structural fragmentation. This paper proposes RPPNet-a two-stage deep learning architecture with variable structural boundaries. It first generates variable-length Rhythm-Pitch Primitive (RPP) sequences, where each RPP encodes note count, rhythm, and contour; then decodes the RPP sequences into concrete notes. The grouping of RPPs is automatically derived from acoustic cues, auditory inertia, and similarity perception based on music psychology. Experiments show that melodies generated by RPPNet are superior in both long-term structure and musicality, with significant improvements across all subjective evaluation dimensions. Ablation studies confirm that the performance gain stems from the structural correctness of the psychological representation, rather than from model capacity. This work offers an interdisciplinary perspective for music generation, integrating music theory, computational modeling, and music psychology.

音乐生成感知建模长程结构心理机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。