arXiv:2604.25498cs.SDcs.AI2026-04中稿 · ISMIR 2026被引 1

用分层结构生成交响乐,可控制和声骨架,音准更和谐。

SymphonyGen: 3D Hierarchical Orchestral Generation with Controllable Harmony Skeleton

论文配图:SymphonyGen: 3D Hierarchical Orchestral Generation with Controllable Harmony Skeleton
图 1 · 摘自论文原文
  • 分三轴解码:小节、声部、音符,降低计算负担
  • 通过和声骨架控制整体结构,减少不协和音
  • 适合作曲初学者与需要可控生成的音乐创作者

生成交响乐需同时处理宏观结构与密集多轨配器,现有符号模型常在可扩展性与可控性间失衡。我们提出SymphonyGen,一种3D分层框架,通过级联解码器分解小节、声部和事件轴,使解码内存远低于传统序列,支持各层级条件输入。采用节拍对齐的多音高和声骨架(可由用户编写、分析或模型生成)作为简谱式条件,实现结构轮廓控制并生成丰富织体。模型通过对抗CLaMP 3音频嵌入的跨模态奖励进行强化学习优化,并引入避不协和采样算法抑制推理中的意外音程冲突。客观评估显示,后训练机制有效降低不协和度,同时保持独立旋律指标;主观测试中,该模型在质量与偏好上均优于基线系统,尤其受普通听众显著青睐。

原文摘要 · Abstract (English)

Generating symphonic music requires simultaneously managing high-level structural form and dense, multi-track orchestration, yet existing symbolic models often struggle with a "complexity-control imbalance" between scalability and steerability. We present SymphonyGen, a 3D hierarchical framework for contemporary orchestral generation, whose cascading decoders decompose the bar, track, and event axes, keeping decoding memory far below flat token streams and enabling conditioning at every structural level. A beat-quantized multi-pitch harmony skeleton, which may be user-written, analyzed, or model-generated, provides "short-score" conditioning, enabling outline control while producing orchestral textures. The model is refined with reinforcement learning against a cross-modal acoustic reward from CLaMP 3 audio embeddings, and a dissonance-averse sampling algorithm suppresses unintended tonal clashes during inference. Objective evaluations show that both post-training mechanisms reduce dissonance while maintaining independent melodic metrics, and in subjective tests SymphonyGen is rated above baseline systems in quality and preference, significantly so among general listeners.

交响乐生成分层建模和声控制音乐生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。