用预训练音乐模型提升节拍追踪准确率
BeatFM: Improving Beat Tracking with Pre-trained Music Foundation Model
- 引入预训练音乐基础模型,融合多维度语义信息
- 在多个基准数据集上达到当前最优节拍追踪效果
- 适合需要高精度节拍分析的音乐智能应用
节拍追踪是音乐信息检索中的经典课题,但现有方法受限于标注数据稀缺,难以泛化到多样音乐风格并精准捕捉复杂节奏结构。为此,我们提出新范式BeatFM,利用预训练音乐基础模型的丰富语义知识提升节拍追踪性能。该模型在多样化音乐数据集上预训练,获得对音乐的强理解能力。为适配节拍追踪任务,设计了即插即用的多维语义聚合模块,包含时序、频率和通道三个并行子模块。大量实验表明,该方法在多个基准数据集上实现了节拍与强拍追踪的最先进性能。
原文摘要 · Abstract (English)
Beat tracking is a widely researched topic in music information retrieval. However, current beat tracking methods face challenges due to the scarcity of labeled data, which limits their ability to generalize across diverse musical styles and accurately capture complex rhythmic structures. To overcome these challenges, we propose a novel beat tracking paradigm BeatFM, which introduces a pre-trained music foundation model and leverages its rich semantic knowledge to improve beat tracking performance. Pre-training on diverse music datasets endows music foundation models with a robust understanding of music, thereby effectively addressing these challenges. To further adapt it for beat tracking, we design a plug-and-play multi-dimensional semantic aggregation module, which is composed of three parallel sub-modules, each focusing on semantic aggregation in the temporal, frequency, and channel domains, respectively. Extensive experiments demonstrate that our method achieves state-of-the-art performance in beat and downbeat tracking across multiple benchmark datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。