arXiv:2502.16936cs.SDcs.AI2025-02ICML被引 14

弱标注音频片段中学习音乐版本匹配,提升段落级识别效果

Supervised contrastive learning from weakly-labeled audio segments for musical version matching

  • 基于片段间距离缩减的弱监督学习方法
  • 新对比损失在段落级任务上实现突破性性能
  • 适用于音频外更多需要细粒度匹配的领域

检测音乐版本(同一乐曲的不同演绎)是一项具有重要应用价值的挑战性任务。现有方法多在整首歌曲级别进行匹配,但多数实际场景需在短片段级别(如20秒)完成匹配。此外,现有方法依赖分类和三元组损失,忽略了近期可能带来显著提升的损失函数。本文提出一种从弱标注片段中学习的方法,并设计了一种改进的对比损失,该损失通过解耦、超参数优化与几何考量进行调整。结合这两项技术,不仅在标准的整首歌级别评估中达到当前最优表现,更在片段级别评估中取得突破性成果。由于所解决的问题具有普遍性,本方法有望推广至音频以外的多个领域。

原文摘要 · Abstract (English)

Detecting musical versions (different renditions of the same piece) is a challenging task with important applications. Because of the ground truth nature, existing approaches match musical versions at the track level (e.g., whole song). However, most applications require to match them at the segment level (e.g., 20s chunks). In addition, existing approaches resort to classification and triplet losses, disregarding more recent losses that could bring meaningful improvements. In this paper, we propose a method to learn from weakly annotated segments, together with a contrastive loss variant that outperforms well-studied alternatives. The former is based on pairwise segment distance reductions, while the latter modifies an existing loss following decoupling, hyper-parameter, and geometric considerations. With these two elements, we do not only achieve state-of-the-art results in the standard track-level evaluation, but we also obtain a breakthrough performance in a segment-level evaluation. We believe that, due to the generality of the challenges addressed here, the proposed methods may find utility in domains beyond audio or musical version matching.

音乐匹配对比学习弱监督片段级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。