arXiv:2510.03025eess.AScs.SD2025-10

通过对比学习提升音乐中人声相似性建模效果

CVSM: Contrastive Vocal Similarity Modeling

  • 采用对比框架,强化相同人声在混音中的相似性
  • 在人声与完整音乐混合物上均超越多个基线模型
  • 无需歌手标签也能达到接近有标签的效果

大规模无标签数据推动了自监督预训练方法的发展。本文提出CVSM(对比人声相似性建模),一种用于音频领域音乐信号表示学习的对比自监督方法,适用于音乐与人声相似性建模。该方法在对比框架下最大化包含相同人声的片段与混音之间的相似性,设计了两种方案:一是利用歌手身份信息构建对比对的有标签协议;二是通过随机抽取人声与伴奏片段人工合成混音,与同段人声配对的无标签方案。我们在客观和主观两个层面评估该方法:客观上通过下游任务线性探测,主观上通过用户研究进行成对比较,结果表明CVSM学习到的表示在人声与音乐相似性建模上表现优异,优于多个基线。尽管有标签预训练整体性能更稳定,但结合真实与人工混音的混合预训练无标签变体,在歌手识别和感知人声相似性方面表现与有标签版本相当。

原文摘要 · Abstract (English)

The availability of large, unlabeled datasets across various domains has contributed to the development of a plethora of methods that learn representations for multiple target (downstream) tasks through self-supervised pre-training. In this work, we introduce CVSM (Contrastive Vocal Similarity Modeling), a contrastive self-supervised procedure for music signal representation learning in the audio domain that can be utilized for musical and vocal similarity modeling. Our method operates under a contrastive framework, maximizing the similarity between vocal excerpts and musical mixtures containing the same vocals; we devise both a label-informed protocol, leveraging artist identity information to sample the contrastive pairs, and a label-agnostic scheme, involving artificial mixture creation from randomly sampled vocal and accompaniment excerpts, which are paired with vocals from the same audio segment. We evaluate our proposed method in measuring vocal similarity both objectively, through linear probing on a suite of appropriate downstream tasks, and subjectively, via conducting a user study consisting of pairwise comparisons between different models in a recommendation-by-query setting. Our results indicate that the representations learned through CVSM are effective in musical and vocal similarity modeling, outperforming numerous baselines across both isolated vocals and complete musical mixtures. Moreover, while the availability of artist identity labels during pre-training leads to overall more consistent performance both in the evaluated downstream tasks and the user study, a label-agnostic CVSM variant incorporating hybrid pre-training with real and artificial mixtures achieves comparable performance to the label-informed one in artist identification and perceived vocal similarity.

自监督学习人声建模对比学习音频表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。