arXiv:2504.07406cs.SDeess.AS2025-04被引 2

构建新数据集并设计音色感知模型,提升吉他音频转录准确性

Towards Generalizability to Tone and Content Variations in the Transcription of Amplifier Rendered Electric Guitar Audio

  • 引入音色嵌入机制,让Transformer模型适应不同音箱配置的音色变化
  • 在多类型音箱配置下,转录准确率显著优于现有方法
  • 适合音乐信号处理、乐器自动转录方向的研究者参考

由于缺乏多样数据集以及放大器、音箱和效果器带来的复杂音色变化,吉他音频转录极具挑战。为此,我们提出EGDB-PG数据集,涵盖多种音箱-音箱组合下的丰富音色特征。同时,我们设计了音色感知变压器(TIT),通过音色嵌入机制利用学习到的表示,增强模型对音色细微差异的适应能力。实验表明,TIT在EGDB-PG上训练后,在多种音箱类型下均优于现有基线,性能提升得益于数据多样性与音色嵌入技术。通过详尽的基准测试与消融实验,我们评估了音色增强、内容增强、音频归一化及音色嵌入对转录性能的影响。本工作突破了数据多样性与音色建模的局限,为未来研究提供了坚实基础。

原文摘要 · Abstract (English)

Transcribing electric guitar recordings is challenging due to the scarcity of diverse datasets and the complex tone-related variations introduced by amplifiers, cabinets, and effect pedals. To address these issues, we introduce EGDB-PG, a novel dataset designed to capture a wide range of tone-related characteristics across various amplifier-cabinet configurations. In addition, we propose the Tone-informed Transformer (TIT), a Transformer-based transcription model enhanced with a tone embedding mechanism that leverages learned representations to improve the model's adaptability to tone-related nuances. Experiments demonstrate that TIT, trained on EGDB-PG, outperforms existing baselines across diverse amplifier types, with transcription accuracy improvements driven by the dataset's diversity and the tone embedding technique. Through detailed benchmarking and ablation studies, we evaluate the impact of tone augmentation, content augmentation, audio normalization, and tone embedding on transcription performance. This work advances electric guitar transcription by overcoming limitations in dataset diversity and tone modeling, providing a robust foundation for future research.

吉他转录音色建模Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。