arXiv:2409.04702cs.SDeess.AS2024-09中稿 · appear in ISMIR 20…被引 20

用新模型同时提升人声分离和旋律转录效果

Mel-RoFormer for Vocal Separation and Vocal Melody Transcription

  • 前端加梅尔带投影,增强多频段特征捕捉能力
  • 分列频率与时间维度建模,提升时频结构理解
  • 先分人声再转旋律,两任务共享基础模型

构建通用深度神经网络以建模音乐音频在音乐信息检索(MIR)中至关重要。由于音乐信号固有的复杂频谱变化,包含旋律、泛音及各类乐器音色,建模难度大。本文提出Mel-RoFormer,一种基于频谱图的模型,包含两项关键设计:前端创新的梅尔带投影模块,增强模型对多频段信息的捕捉能力;以及交错式位置编码变压器(RoPE Transformers),将频率与时间维度作为独立序列显式建模。该模型被应用于两个核心MIR任务:人声分离与人声旋律转录,分别旨在从混音中分离出人声,并转录其主旋律。尽管两者均聚焦于人声信号,但优化目标不同。因此采用两步法:先训练人声分离模型,再以此为基础微调用于旋律转录。在基准数据集上的大量实验表明,该模型在人声分离与旋律转录任务上均达到当前最优性能,验证了Mel-RoFormer在建模复杂音乐信号方面的有效性与通用性。

原文摘要 · Abstract (English)

Developing a versatile deep neural network to model music audio is crucial in MIR. This task is challenging due to the intricate spectral variations inherent in music signals, which convey melody, harmonics, and timbres of diverse instruments. In this paper, we introduce Mel-RoFormer, a spectrogram-based model featuring two key designs: a novel Mel-band Projection module at the front-end to enhance the model's capability to capture informative features across multiple frequency bands, and interleaved RoPE Transformers to explicitly model the frequency and time dimensions as two separate sequences. We apply Mel-RoFormer to tackle two essential MIR tasks: vocal separation and vocal melody transcription, aimed at isolating singing voices from audio mixtures and transcribing their lead melodies, respectively. Despite their shared focus on singing signals, these tasks possess distinct optimization objectives. Instead of training a unified model, we adopt a two-step approach. Initially, we train a vocal separation model, which subsequently serves as a foundation model for fine-tuning for vocal melody transcription. Through extensive experiments conducted on benchmark datasets, we showcase that our models achieve state-of-the-art performance in both vocal separation and melody transcription tasks, underscoring the efficacy and versatility of Mel-RoFormer in modeling complex music audio signals.

人声分离旋律转录频谱建模Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。