arXiv:2510.25560cs.SDeess.AS2025-10被引 1

用多假设筛选提升音乐节拍追踪的自监督学习效果

Controlling Contrastive Self-Supervised Learning with Knowledge-Driven Multiple Hypothesis: Application to Beat Tracking

  • 基于领域知识筛选多种可能的正样本假设
  • 在标准数据集上优于现有方法,提升节拍追踪精度
  • 适合需要融合专家知识的音乐表示学习场景

数据和问题约束中的模糊性可能导致机器学习任务存在多种合理结果。以节拍与强拍追踪为例,不同听者可能采用不同的节奏解读,均不必然错误。为此,我们提出一种对比自监督预训练方法,利用数据中关于正样本的多个假设进行训练。模型学习与多种假设兼容的表示,并通过基于知识的评分函数选择最合理的假设。在标注数据上微调后,该模型在标准基准测试中表现优于现有方法,证明了将领域知识与多假设选择相结合在音乐表征学习中的优势。

原文摘要 · Abstract (English)

Ambiguities in data and problem constraints can lead to diverse, equally plausible outcomes for a machine learning task. In beat and downbeat tracking, for instance, different listeners may adopt various rhythmic interpretations, none of which would necessarily be incorrect. To address this, we propose a contrastive self-supervised pre-training approach that leverages multiple hypotheses about possible positive samples in the data. Our model is trained to learn representations compatible with different such hypotheses, which are selected with a knowledge-based scoring function to retain the most plausible ones. When fine-tuned on labeled data, our model outperforms existing methods on standard benchmarks, showcasing the advantages of integrating domain knowledge with multi-hypothesis selection in music representation learning in particular.

音乐理解自监督学习多假设

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。