arXiv:2509.23759cs.SDcs.LG2025-09被引 3

首个联合识别小提琴演奏技巧与音高时长的模型

VioPTT: Violin Technique-Aware Transcription from Synthetic Data Augmentation

  • 轻量级级联模型,同时预测音高、时长与演奏技巧
  • 在真实小提琴录音上达到顶尖转录性能,泛化能力强
  • 自建高质量合成数据集,避免人工标注瓶颈

尽管音乐信息检索中的自动音乐转录已较为成熟,但现有模型多仅关注音高和时间信息,忽略了表达性与乐器特异性细节。以小提琴为例,其演奏技巧直接影响音色表现力与情感张力。本文提出VioPTT(小提琴演奏技巧感知转录)模型,可直接输出小提琴演奏技巧,同时预测音高起止时间。我们还发布了MOSA-VPT——一个高质量、新颖的小提琴演奏技巧合成数据集,以规避人工标注需求。基于该数据集训练的模型,在真实小提琴记谱数据上表现出优异泛化能力,并达到当前最佳转录性能。据我们所知,VioPTT是首个在统一框架中联合实现小提琴转录与演奏技巧预测的模型。

原文摘要 · Abstract (English)

While automatic music transcription is well-established in music information retrieval, most models are limited to transcribing pitch and timing information from audio, and thus omit crucial expressive and instrument-specific nuances. One example is playing technique on the violin, which affords its distinct palette of timbres for maximal emotional impact. Here, we propose VioPTT (Violin Playing Technique-aware Transcription), a lightweight cascade model that directly transcribes violin playing technique in addition to pitch onset and offset. Furthermore, we release MOSA-VPT, a novel, high-quality synthetic violin playing technique dataset to circumvent the need for manually labeled annotations. Leveraging this dataset, our model demonstrated strong generalization to real-world note-level violin technique recordings in addition to achieving state-of-the-art transcription performance. To our knowledge, VioPTT is the first to jointly combine violin transcription and playing technique prediction within a unified framework.

小提琴音乐转录合成数据技巧识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。