首个含音频、乐谱与文本标签的钢琴数据集,助力音乐信息检索研究。
PIAST: A Multimodal Piano Dataset with Audio, Symbolic and Text
- 构建基于钢琴语义标签的多模态数据集,涵盖音视频、乐谱与人工标注文本。
- 收录9673首曲目,2023首经专家标注,支持音频与MIDI的音乐标签与检索任务。
- 适合音乐信息检索、跨模态学习及钢琴音乐分析方向的研究者使用。
尽管钢琴音乐已成为音乐信息检索(MIR)的重要研究方向,但缺乏带有文本标签的钢琴独奏数据集。为此,我们提出PIAST(PIano dataset with Audio, Symbolic, and Text),一个包含音频、符号表示与文本标签的钢琴音乐数据集。基于钢琴特定的语义标签体系,从YouTube收集了9,673首曲目,并由音乐专家对其中2,023首进行人工标注,形成PIAST-YT和PIAST-AT两个子集。两个子集均包含音频、文本标签、标签注释以及通过先进钢琴转录与节拍追踪模型生成的MIDI乐谱。我们在此多模态数据集上开展了音乐标签与检索任务,报告基线性能,以展示其在MIR研究中的潜力。
原文摘要 · Abstract (English)
While piano music has become a significant area of study in Music Information Retrieval (MIR), there is a notable lack of datasets for piano solo music with text labels. To address this gap, we present PIAST (PIano dataset with Audio, Symbolic, and Text), a piano music dataset. Utilizing a piano-specific taxonomy of semantic tags, we collected 9,673 tracks from YouTube and added human annotations for 2,023 tracks by music experts, resulting in two subsets: PIAST-YT and PIAST-AT. Both include audio, text, tag annotations, and transcribed MIDI utilizing state-of-the-art piano transcription and beat tracking models. Among many possible tasks with the multi-modal dataset, we conduct music tagging and retrieval using both audio and MIDI data and report baseline performances to demonstrate its potential as a valuable resource for MIR research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。