用GAN生成逼真手写乐谱符号,解决历史乐谱数据少的问题
GAN-based Content-Conditioned Generation of Handwritten Musical Symbols
- 基于音乐符号级GAN生成手写风格乐符
- 生成符号视觉真实度高,可有效提升识别模型性能
- 适合音乐信息学、OMR研究者及数据匮乏场景
光学乐谱识别(OMR)因真实标注数据稀缺而受限,尤其在处理手写历史乐谱时。类似地,在手写文本识别领域已证明,通过图像生成技术合成样本可提升识别模型性能。本研究采用音乐符号级生成对抗网络(GAN),生成逼真手写风格的乐符,并利用Smashcima排版软件将其组合成完整乐谱。系统评估了生成样本的视觉保真度,结果表明生成符号具有高度真实性,标志着合成乐谱生成的重要进展。
原文摘要 · Abstract (English)
The field of Optical Music Recognition (OMR) is currently hindered by the scarcity of real annotated data, particularly when dealing with handwritten historical musical scores. In similar fields, such as Handwritten Text Recognition, it was proven that synthetic examples produced with image generation techniques could help to train better-performing recognition architectures. This study explores the generation of realistic, handwritten-looking scores by implementing a music symbol-level Generative Adversarial Network (GAN) and assembling its output into a full score using the Smashcima engraving software. We have systematically evaluated the visual fidelity of these generated samples, concluding that the generated symbols exhibit a high degree of realism, marking significant progress in synthetic score generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。