arXiv:2503.03232cs.SDeess.AS2025-03

通过多轨音频精标注与注意力机制,实现乐器精准识别。

Lead Instrument Detection from Multitrack Music

  • 构建多轨音频标注数据集,结合自监督学习与逐轨注意力分类器。
  • 在未见乐器和跨域测试中表现更优,准确率显著超越传统模型。
  • 适合音乐内容分析、智能编曲等音频处理研究者参考。

现有主流方法多基于混合音频进行分析,仅能实现粗粒度分类且泛化能力差。本文提出一种新型多轨音乐音频中的主奏乐器检测方法,通过构建精心标注的数据集,并设计融合自监督学习与逐轨帧级注意力机制的分类框架。该注意力机制根据听觉重要性动态提取并聚合各音轨特征,实现对多种乐器类型及组合的精确检测。结合音轨分类与排列增强策略,所提模型在未见乐器和跨域测试中均表现稳健,显著优于SVM与CRNN等现有模型。研究为多轨音乐音频内容分析提供了新思路。

原文摘要 · Abstract (English)

Prior approaches to lead instrument detection primarily analyze mixture audio, limited to coarse classifications and lacking generalization ability. This paper presents a novel approach to lead instrument detection in multitrack music audio by crafting expertly annotated datasets and designing a novel framework that integrates a self-supervised learning model with a track-wise, frame-level attention-based classifier. This attention mechanism dynamically extracts and aggregates track-specific features based on their auditory importance, enabling precise detection across varied instrument types and combinations. Enhanced by track classification and permutation augmentation, our model substantially outperforms existing SVM and CRNN models, showing robustness on unseen instruments and out-of-domain testing. We believe our exploration provides valuable insights for future research on audio content analysis in multitrack music settings.

乐器检测多轨音频注意力机制自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。