用原始MIDI信息改进音乐变换器的结构表示,效果更优且省数据标注成本
Do we need more complex representations for structure? A comparison of note duration representation for Music Transformers
- 用原始MIDI数据替代标注结构,简化输入表示
- 微调表示方法后,结构相似性指标提升显著
- 适合关注低成本音乐生成的开发者和研究者
近年来,深度学习在创意计算领域取得显著成果。音乐生成中,基于Transformer的模型表现突出,但通常依赖人工标注的结构信息。本文探究:仅使用未标注的MIDI数据时,现成的Music Transformer能否在结构相似性指标上表现良好。结果表明,对最常见表示方式稍作调整,即可带来小而显著的性能提升。研究主张,探索更优的无标注音乐表示,比大规模人工标注数据更具成本效益。
原文摘要 · Abstract (English)
In recent years, deep learning has achieved formidable results in creative computing. When it comes to music, one viable model for music generation are Transformer based models. However, while transformers models are popular for music generation, they often rely on annotated structural information. In this work, we inquire if the off-the-shelf Music Transformer models perform just as well on structural similarity metrics using only unannotated MIDI information. We show that a slight tweak to the most common representation yields small but significant improvements. We also advocate that searching for better unannotated musical representations is more cost-effective than producing large amounts of curated and annotated data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。