arXiv:2410.17209cs.SDcs.CL2024-10

用Whisper框架实现音频转乐谱,支持旋律和和弦识别

Audio-to-Score Conversion Model Based on Whisper methodology

  • 基于Whisper的Transformer模型,将音频转为ABC记谱法
  • 自研'奥菲斯乐谱'系统,提升训练数据多样性和模型精度
  • 适合音乐信息处理研究者与爱好者使用

本论文构建了一个基于Whisper的Transformer模型,可从音乐音频中提取旋律与和弦,并以ABC记谱法记录。针对ABC记谱法设计了完整的数据处理流程,包括清洗、格式化与转换,并引入变异机制以增强训练数据的多样性与质量。论文创新性地提出“奥菲斯乐谱”这一自定义记谱体系,将音乐信息转化为标记符,构建专属词典并训练对应分词器。实验表明,相比传统算法,该模型在准确率与性能上均有显著提升。该工作不仅为音乐爱好者提供便捷的音频转乐谱工具,也为音乐信息处理领域的研究提供了新思路与实用工具。

原文摘要 · Abstract (English)

This thesis develops a Transformer model based on Whisper, which extracts melodies and chords from music audio and records them into ABC notation. A comprehensive data processing workflow is customized for ABC notation, including data cleansing, formatting, and conversion, and a mutation mechanism is implemented to increase the diversity and quality of training data. This thesis innovatively introduces the "Orpheus' Score", a custom notation system that converts music information into tokens, designs a custom vocabulary library, and trains a corresponding custom tokenizer. Experiments show that compared to traditional algorithms, the model has significantly improved accuracy and performance. While providing a convenient audio-to-score tool for music enthusiasts, this work also provides new ideas and tools for research in music information processing.

音频转乐谱WhisperABC记谱法音乐信息处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。