arXiv:2608.03419cs.SDcs.AI2026-08中稿 · the 27th Internati…

首个完整视觉钢琴转录系统,精准识别音符起止与力度

Multi-Task Multi-Frame Visual Piano Transcription

论文配图:Multi-Task Multi-Frame Visual Piano Transcription
图 1 · 摘自论文原文
  • 共享时序主干+多任务头,逐帧监督提升整体性能
  • 在PianoVAM和R3数据集上刷新音符起止与力度预测纪录
  • 适合音乐视频分析、自动记谱等需要高精度音符信息的场景

基于音频的钢琴转录在音符起始、音高和力度上表现良好,但延音踏板会使声音在按键释放后仍持续,导致音频系统预测的是踏板延长的结束时间而非物理按键释放时间。现有视觉钢琴转录(VPT)系统主要关注短视频窗口内的起始点检测,对终止点的准确性严重不足,且未报告音符级力度信息。为此,我们提出V2N(Video to Notes),首个完整的视觉钢琴转录系统:采用共享时序主干网络,搭配针对起始、终止、键位保持和力度的任务特定头,通过逐帧监督而非仅在窗口中心监督进行联合训练。消融实验表明,多任务监督提升了终止点与力度预测能力,同时改善了起始点准确率;更长的时序上下文进一步提升性能。V2N在PianoVAM和R3数据集上达到新最优结果。

原文摘要 · Abstract (English)

Audio-based piano transcription performs well on onset, pitch, and velocity, but the sustain pedal lets sound persist long after key release, so audio systems predict pedal-extended offsets rather than physical key release. Yet existing Visual Piano Transcription (VPT) systems focus on onset detection from short video windows, offset accuracy lags onset by a wide margin, and note-level velocity has not been reported. To address these gaps, we present V2N (Video to Notes), the first complete VPT system: a shared temporal backbone feeds task-specific heads for onset, offset, key hold, and velocity, jointly trained with per-frame supervision rather than only at the window center. Ablations show that multi-task supervision enables offset and velocity prediction while improving onset accuracy; longer temporal context yields further improvements. V2N sets new state-of-the-art results on PianoVAM and R3.

视觉转录钢琴识别多任务学习时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。