arXiv:2601.14157cs.SDcs.AI2026-01

构建21k条音乐概念数据集,提升音乐模型可解释性。

ConceptCaps: a Distilled Concept Dataset for Interpretability in Music Models

  • 分步构建:先学属性共现,再生成描述,最后合成音频。
  • 在CLAP和TCAV测试中,概念探针准确捕捉音乐模式。
  • 适合研究音乐生成、可解释性或风格控制的学者。

基于概念的可解释性方法(如TCAV)需要为每个概念提供清晰、分离的正负样本,但现有音乐数据集标签稀疏、嘈杂或定义模糊。我们提出ConceptCaps,一个包含21,000个音乐-文本-标签三元组的数据集,基于200个属性的分类体系标注。我们的流程将语义建模与文本生成分离:使用变分自编码器(VAE)学习合理的属性共现模式,微调大语言模型(LLM)将属性列表转化为专业描述,再由MusicGen生成对应音频。这种分离提升了生成内容的连贯性与可控性,优于端到端方法。通过音频-文本对齐(CLAP)、语言质量指标(BERTScore、MAUVE)以及TCAV分析验证,概念探针能有效识别出具有音乐意义的模式。数据集与代码已公开。

原文摘要 · Abstract (English)

Concept-based interpretability methods like TCAV require clean, well-separated positive and negative examples for each concept. Existing music datasets lack this structure: tags are sparse, noisy, or ill-defined. We introduce ConceptCaps, a dataset of 21k music-caption-tags triplets with explicit labels from a 200-attribute taxonomy. Our pipeline separates semantic modeling from text generation: a VAE learns plausible attribute co-occurrence patterns, a fine-tuned LLM converts attribute lists into professional descriptions, and MusicGen synthesizes corresponding audio. This separation improves coherence and controllability over end-to-end approaches. We validate the dataset through audio-text alignment (CLAP), linguistic quality metrics (BERTScore, MAUVE), and TCAV analysis confirming that concept probes recover musically meaningful patterns. Dataset and code are available online.

音乐生成可解释性数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。