arXiv:2511.07268cs.SDeess.AS2025-11

对比不同模型与数据对钢琴乐谱生成质量的影响。

Generating Piano Music with Transformers: A Comparative Study of Scale, Data, and Metrics

  • 系统比较数据集、模型结构与训练策略对音乐生成的影响。
  • 950M参数模型在8万首多风格乐曲上训练,生成效果接近真人创作。
  • 验证了多种评估指标与人类听觉判断的相关性,指导模型优化。

尽管近年来已提出多种用于符号化音乐生成的Transformer模型,但关于具体设计选择如何影响生成音乐质量的综合性研究仍较少。本文系统比较了符号化钢琴音乐生成任务中不同的数据集、模型架构、模型规模和训练策略。为支持模型开发与评估,我们考察了一系列量化指标,并分析其与通过听觉实验收集的人类判断之间的相关性。我们表现最佳的模型是一个在8万首来自多元音乐风格的MIDI文件上训练的950M参数Transformer,在类似图灵测试的听觉调查中,其生成结果常被误认为是人类创作的作品。

原文摘要 · Abstract (English)

Although a variety of transformers have been proposed for symbolic music generation in recent years, there is still little comprehensive study on how specific design choices affect the quality of the generated music. In this work, we systematically compare different datasets, model architectures, model sizes, and training strategies for the task of symbolic piano music generation. To support model development and evaluation, we examine a range of quantitative metrics and analyze how well they correlate with human judgment collected through listening studies. Our best-performing model, a 950M-parameter transformer trained on 80K MIDI files from diverse genres, produces outputs that are often rated as human-composed in a Turing-style listening survey.

音乐生成Transformer评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。