首个系统分段的乐谱图像数据集,用于深度学习研究。
VisionScores -- A system-segmented image score dataset for deep learning tasks
- 按作曲家和曲式类型构建双场景乐谱图像数据集。
- 共24.8万张128×512灰度图,含14k Sonatina与10.8k李斯特作品。
- 提供分段图像、原始页面及元数据,支持多维度分析。
VisionScores 是首个系统分段的乐谱图像数据集,旨在为机器学习与深度学习任务提供结构丰富、信息密度高的图像数据。数据集聚焦于双手钢琴曲,考虑图形相似性与创作模式,因创作过程高度依赖乐器特性。包含两个场景:其一为14,000个样本,涵盖不同作曲家但同为小奏鸣曲(Sonatinas);其二为10,800个样本,同一作曲家(李斯特)的不同曲式类型。全部24.8万个样本均以128×512像素的灰度jpg格式呈现。除格式化图像外,还提供系统的顺序信息、曲目元数据,以及未分段的全页乐谱和预处理图像,便于后续分析。
原文摘要 · Abstract (English)
VisionScores presents a novel proposal being the first system-segmented image score dataset, aiming to offer structure-rich, high information-density images for machine and deep learning tasks. Delimited to two-handed piano pieces, it was built to consider not only certain graphic similarity but also composition patterns, as this creative process is highly instrument-dependent. It provides two scenarios in relation to composer and composition type. The first, formed by 14k samples, considers works from different authors but the same composition type, specifically, Sonatinas. The latter, consisting of 10.8K samples, presents the opposite case, various composition types from the same author, being the one selected Franz Liszt. All of the 24.8k samples are formatted as grayscale jpg images of $128 \times 512$ pixels. VisionScores supplies the users not only the formatted samples but the systems' order and pieces' metadata. Moreover, unsegmented full-page scores and the pre-formatted images are included for further analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。