arXiv:2411.09037cs.CV2024-11IJCAI被引 1

用视觉变压器实现钢琴演奏的精准符号化转录。

Pay Attention to the Keys: Visual Piano Transcription Using Transformers

  • 基于视觉变压器,从键盘视频中识别音符起始时刻。
  • 在R3数据集上同时提升起始与偏移预测准确率。
  • 首次探索音符偏移预测,适合音乐视觉分析研究者。

视觉钢琴转录(VPT)是从纯视觉信息(如俯视钢琴键盘视频)中获取钢琴演奏的符号表示。本文提出一种基于视觉变压器(ViT)的VPT系统,超越了以往基于卷积神经网络(CNN)的方法。系统在新提出的R3数据集上进行训练,该数据集包含约31小时同步的视频与MIDI录音。我们还引入了一种预测音符偏移的新方法,此前未在该领域被探索。实验表明,本系统在PianoYT数据集上优于现有最优的起始时刻预测表现,在R3数据集上对起始和偏移预测均取得领先结果。

原文摘要 · Abstract (English)

Visual piano transcription (VPT) is the task of obtaining a symbolic representation of a piano performance from visual information alone (e.g., from a top-down video of the piano keyboard). In this work we propose a VPT system based on the vision transformer (ViT), which surpasses previous methods based on convolutional neural networks (CNNs). Our system is trained on the newly introduced R3 dataset, consisting of ca.~31 hours of synchronized video and MIDI recordings of piano performances. We additionally introduce an approach to predict note offsets, which has not been previously explored in this context. We show that our system outperforms the state-of-the-art on the PianoYT dataset for onset prediction and on the R3 dataset for both onsets and offsets.

视觉转录变压器音乐分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。