arXiv:2412.06703cs.SDcs.AI2024-12被引 2

用深度学习将混音音频分离为乐器声部并转成乐谱

Source Separation & Automatic Transcription for Music

  • 通过频谱掩码与神经网络分离音频中的不同乐器
  • 能从wav文件生成对应MIDI和可读乐谱
  • 适合音乐制作人与自动记谱研究者

音源分离是从多重声音混合中提取单一声音的过程,应用广泛,涵盖语音增强、歌词转录及音乐数字音频制作。自动音乐记谱(AMT)则是将原始音乐音频转换为乐手可读的五线谱。过去这些任务受限于噪声大、训练时间长及版权导致的数据难获取。近年来深度学习带来新方法,可实现低失真音轨分离与音频到乐谱的生成。本文结合频谱掩码、深度神经网络与MuseScore API,构建端到端流程:输入wav格式音乐混合信号后,分离出各乐器声部,生成对应MIDI文件,并转为每件乐器的乐谱。

原文摘要 · Abstract (English)

Source separation is the process of isolating individual sounds in an auditory mixture of multiple sounds [1], and has a variety of applications ranging from speech enhancement and lyric transcription [2] to digital audio production for music. Furthermore, Automatic Music Transcription (AMT) is the process of converting raw music audio into sheet music that musicians can read [3]. Historically, these tasks have faced challenges such as significant audio noise, long training times, and lack of free-use data due to copyright restrictions. However, recent developments in deep learning have brought new promising approaches to building low-distortion stems and generating sheet music from audio signals [4]. Using spectrogram masking, deep neural networks, and the MuseScore API, we attempt to create an end-to-end pipeline that allows for an initial music audio mixture (e.g...wav file) to be separated into instrument stems, converted into MIDI files, and transcribed into sheet music for each component instrument.

音源分离音乐记谱深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。