arXiv:2508.07987cs.SDcs.CL2025-08中稿 · the 6th Conference…

用程序生成吉他指弹数据,解决标注样本少的问题

Exploring Procedural Data Generation for Automatic Acoustic Guitar Fingerpicking Transcription

  • 通过四步流程合成逼真吉他音频数据
  • 合成数据训练模型可达到合理识别准确率
  • 少量真实数据微调后效果优于纯真实数据训练

由于缺乏标注训练数据及音乐录音的法律限制,自动转录原声吉他指弹演奏仍具挑战。本文探索一种基于程序生成的数据流水线,替代真实音频用于训练转录模型。方法包含四个阶段:基于知识的指弹谱创作、MIDI演奏渲染、基于改进卡拉普斯-史龙算法的物理建模,以及包含混响和失真在内的音频增强。我们在真实与合成数据上训练并评估了一个基于CRNN的音符追踪模型,结果表明程序生成数据可用于实现合理的音符追踪效果。使用少量真实数据微调后,模型性能超越仅在真实数据上训练的模型。这些结果凸显了程序生成音频在数据稀缺的音乐信息检索任务中的潜力。

原文摘要 · Abstract (English)

Automatic transcription of acoustic guitar fingerpicking performances remains a challenging task due to the scarcity of labeled training data and legal constraints connected with musical recordings. This work investigates a procedural data generation pipeline as an alternative to real audio recordings for training transcription models. Our approach synthesizes training data through four stages: knowledge-based fingerpicking tablature composition, MIDI performance rendering, physical modeling using an extended Karplus-Strong algorithm, and audio augmentation including reverb and distortion. We train and evaluate a CRNN-based note-tracking model on both real and synthetic datasets, demonstrating that procedural data can be used to achieve reasonable note-tracking results. Finetuning with a small amount of real data further enhances transcription accuracy, improving over models trained exclusively on real recordings. These results highlight the potential of procedurally generated audio for data-scarce music information retrieval tasks.

音频生成音乐转录数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。