用扩散模型生成可弹奏的吉他谱,兼顾音准与指法合理性。
Playability-Aware Audio-to-Tablature Guitar Transcription via Diffusion Models

- 通过连续潜空间建模离散弦品位置,实现更自然的指法生成。
- 引入五种辅助损失,在两个数据集上均超越基线模型。
- 适合音乐科技、自动配乐和吉他教学系统开发者使用。
吉他谱转录不仅需要准确检测音高,还需为每个音符分配具体的弦品位置,因为同一音高可在指板上多个位置演奏。现有方法将其视为标准分类问题,忽略了音乐和物理层面限制可演奏指法序列的因素。我们提出 Noise2Fret,一种基于扩散模型的音频到吉他谱转录方法,通过条件于频谱和音频特征的连续潜表示生成吉他谱,目标为离散的弦品组合。为弥合音高准确率与实际可演奏性之间的差距,我们在训练目标中引入五种辅助损失:音级距离、位置距离、五度圈距离、弦相似性及手跨度可行性。在 GuitarSet 与 GOAT 数据集上的实验表明,该模型在性能上优于基线方法,同时计算效率更高,且辅助损失带来一致性的提升。
原文摘要 · Abstract (English)
Guitar tablature transcription requires not only accurate pitch detection but also assigning each note to a specific string-fret position, as the same pitch can be played at multiple fretboard positions. Existing approaches treat this as a standard classification problem, ignoring the musical and physical constraints that govern playable fingering sequences. We propose Noise2Fret, a diffusion model for audio-to-tablature transcription that generates tablature through a continuous latent representation of discrete fret and string targets, conditioned on spectral and audio features. To bridge the gap between pitch accuracy and physical playability, we introduce five auxiliary losses encoding Pitch-Class Distance, Positional Distance, Circle-of-Fifths Distance, String Similarity, and Hand-Span Feasibility directly into the training objective. Experiments on GuitarSet and GOAT datasets demonstrate that the model outperforms baselines while remaining computationally more efficient, and that the auxiliary losses yield consistent gains over the standard training objective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。