arXiv:2509.08800cs.SDcs.AI2025-09中稿 · the 26th Internati…被引 8

构建首个包含多模态数据的钢琴演奏数据集,助力音乐信息检索研究。

PianoVAM: A Multimodal Piano Performance Dataset

  • 采集业余钢琴家日常练习的视频、音频、MIDI等多模态数据。
  • 通过手部姿态模型与半自动标注法生成手指标记与指法标签。
  • 支持音视频联合转录等任务,适合音乐智能与人机交互研究者。

音乐表演的多模态特性推动了音乐信息检索(MIR)领域对音频之外数据的兴趣。本文提出PianoVAM,一个全面的钢琴演奏数据集,包含视频、音频、MIDI、手部关键点、指法标签和丰富元数据。数据通过Disklavier钢琴录制,捕捉业余钢琴家日常练习中的音频与MIDI信号,并同步获取俯视视角视频,覆盖真实多样演奏场景。手部关键点与指法标签通过预训练手部姿态估计模型及半自动指法标注算法提取。文中讨论了数据收集与跨模态对齐过程中的挑战。此外,基于视频中提取的手部关键点描述了指法标注方法。最后,展示了使用PianoVAM进行纯音频与音视频联合钢琴转录的基准测试结果,并探讨了其他潜在应用。

原文摘要 · Abstract (English)

The multimodal nature of music performance has driven increasing interest in data beyond the audio domain within the music information retrieval (MIR) community. This paper introduces PianoVAM, a comprehensive piano performance dataset that includes videos, audio, MIDI, hand landmarks, fingering labels, and rich metadata. The dataset was recorded using a Disklavier piano, capturing audio and MIDI from amateur pianists during their daily practice sessions, alongside synchronized top-view videos in realistic and varied performance conditions. Hand landmarks and fingering labels were extracted using a pretrained hand pose estimation model and a semi-automated fingering annotation algorithm. We discuss the challenges encountered during data collection and the alignment process across different modalities. Additionally, we describe our fingering annotation method based on hand landmarks extracted from videos. Finally, we present benchmarking results for both audio-only and audio-visual piano transcription using the PianoVAM dataset and discuss additional potential applications.

多模态钢琴演奏数据集手部姿态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。