arXiv:2509.19412cs.GRcs.AI2025-09中稿 · the International …

用图神经网络统一解决钢琴乐谱自动排版难题

EngravingGNN: A Hybrid Graph Neural Network for End-to-End Piano Score Engraving

  • 构建多任务图神经网络,联合预测音符连接、谱表分配等7个排版要素
  • 在日系流行与浪漫主义双数据集上,各项指标均优于单一任务模型
  • 适合需要端到端乐谱生成的音乐信息学与数字出版研究者

本文聚焦于自动乐谱排版,即从符号化音乐内容生成可读性强的乐谱。这一环节对所有涉及人类演奏的应用至关重要,但在符号音乐处理领域仍属未充分探索方向。本文将问题形式化为一系列相互关联的子任务,提出针对钢琴音乐和量化符号输入的统一图神经网络(GNN)框架。方法采用多任务GNN,联合预测音部连接、谱表分配、音高记法、调号、符干方向、八度移动及谱号等7项关键元素,并通过专用后处理流程生成可打印的MusicXML/MEI输出。在两个多样化的钢琴数据集(J-Pop与DCML Romantic)上的综合评估表明,该统一模型在所有子任务上均取得良好准确率,优于仅专注于特定任务的现有系统。结果表明,在多任务设置下,共享的GNN编码器搭配轻量级任务专用解码器,是实现自动乐谱排版的可扩展且高效方案。

原文摘要 · Abstract (English)

This paper focuses on automatic music engraving, i.e., the creation of a humanly-readable musical score from musical content. This step is fundamental for all applications that include a human player, but it remains a mostly unexplored topic in symbolic music processing. In this work, we formalize the problem as a collection of interdependent subtasks, and propose a unified graph neural network (GNN) framework that targets the case of piano music and quantized symbolic input. Our method employs a multi-task GNN to jointly predict voice connections, staff assignments, pitch spelling, key signature, stem direction, octave shifts, and clef signs. A dedicated postprocessing pipeline generates print-ready MusicXML/MEI outputs. Comprehensive evaluation on two diverse piano corpora (J-Pop and DCML Romantic) demonstrates that our unified model achieves good accuracy across all subtasks, compared to existing systems that only specialize in specific subtasks. These results indicate that a shared GNN encoder with lightweight task-specific decoders in a multi-task setting offers a scalable and effective solution for automatic music engraving.

乐谱生成图神经网络符号音乐多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。