arXiv:2411.04152eess.AScs.SD2024-11被引 2

无监督训练模型,少量标注数据即可实现精准节拍追踪。

A Contrastive Self-Supervised Learning scheme for beat tracking amenable to few-shot learning

  • 用对比学习让模型区分不同节拍间隔的音频片段。
  • 在仅几例标注数据下,性能媲美传统方法。
  • 适合标注数据稀缺的音乐分析场景。

本文提出一种新型自监督学习方案,用于训练节奏分析系统,并应用于少样本节拍追踪任务。受对比预测编码启发,设计了一个对数梅尔频谱图变换器编码器,通过对比假设节拍间隔内与非间隔内的音频观察来训练。无需真实节拍位置或速度信息,利用主导局部脉冲函数的局部极大值作为音符起点(Tatum)的代理,定义候选锚点、正样本(距离为2的幂次)和负样本(其余时间位置)。在未标注的FMA、MTT和MTG-Jamendo数据集上预训练后,该模型可在仅少量标注样本的情况下微调,获得具有竞争力的节拍追踪性能。

原文摘要 · Abstract (English)

In this paper, we propose a novel Self-Supervised-Learning scheme to train rhythm analysis systems and instantiate it for few-shot beat tracking. Taking inspiration from the Contrastive Predictive Coding paradigm, we propose to train a Log-Mel-Spectrogram Transformer encoder to contrast observations at times separated by hypothesized beat intervals from those that are not. We do this without the knowledge of ground-truth tempo or beat positions, as we rely on the local maxima of a Predominant Local Pulse function, considered as a proxy for Tatum positions, to define candidate anchors, candidate positives (located at a distance of a power of two from the anchor) and negatives (remaining time positions). We show that a model pre-trained using this approach on the unlabeled FMA, MTT and MTG-Jamendo datasets can successfully be fine-tuned in the few-shot regime, i.e. with just a few annotated examples to get a competitive beat-tracking performance.

自监督学习节拍追踪少样本学习音频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。