arXiv:2608.25244cs.SD2026-08中稿 · ISMIR 2026

用专业乐评补全音乐模型训练数据,提升检索与分类效果

AllMusicCaps: Album Reviews as Complementary Supervision for Music CLAP

  • 用AI提取乐评中的描述性语句,生成24.5万条可用训练数据
  • 在复杂查询上检索准确率显著提升,优于现有数据集
  • 新方法适合做音乐理解、跨模态检索等任务的研究者

近期开放的文本-音频对比模型(如CLAP)通常使用大语言模型生成的标签或网络搜索结果作为标题,虽准确但表达狭窄。本文探索了人类撰写的专辑评论作为补充监督信号,特别是来自AllMusic的专业乐评——它们规模可观,包含叙事线索、评价性形容词和场景刻画,是其他数据源缺乏的。由于原始乐评噪声过大,无法直接作为标题使用,我们通过大语言模型预处理流程,从245,346条乐评中提取描述性音乐语句,并重写为可训练的标题数据。实验表明,使用乐评监督在人工标题基准(Song Describer)上取得最大检索提升,尤其对其他标题数据集未覆盖的复杂查询表现优异。此外,我们重新审视训练策略,发现SigReg正则化(使嵌入空间呈各向同性高斯分布)能提升多层感知机探测任务及文本到音乐检索的表现。最终模型在文本到音乐检索、零样本分类和多数探测任务上超越现有公开的CLAP类基线。我们公开了基于乐评生成的标题数据集和模型权重,以支持后续研究。

原文摘要 · Abstract (English)

Recent open text-audio contrastive models (CLAPs) are typically trained with LLM-generated captions derived from tag datasets or web search results, which tend to be accurate but expressively narrow. As a complementary source, we explore human-written album reviews, specifically expert reviews from AllMusic: they exist at scale and carry narrative cues, evaluative adjectives, and scene framing that other sources lack. Since raw reviews are too noisy for direct use as captions, we first build a caption corpus with 24,5346 samples via an LLM preprocessing pipeline that identifies descriptive musical quotes and rewrites them into training-ready captions. We find that album review supervision yields the largest retrieval gains on a human-written caption benchmark (Song Describer), particularly for complex queries that other existing caption datasets leave uncovered. In addition, we revisit the training recipe and show that SigReg regularization, which encourages an isotropic Gaussian distribution in the embedding space, improves MLP probing across classification tasks, as well as text-to-music retrieval. The resulting model outperforms open CLAP-style baselines on text-to-music retrieval, zero-shot classification, and most MLP probing tasks. We release the review-derived caption dataset and model weights to support future research.

音乐理解跨模态乐评数据CLAP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。