arXiv:2606.06615cs.SDcs.AI2026-06ACL

让音乐检索精准匹配慢快、调性等细节,效果提升超70%。

FIGMA: Towards FIne-Grained Music retrievAl

论文配图:FIGMA: Towards FIne-Grained Music retrievAl
图 1 · 摘自论文原文
  • 设计多视角对比架构,同时对齐音频与文本的全局和帧级特征
  • 在38万组音乐-描述对上训练,测试集达1万条,含多种音乐属性标注
  • 首次系统构建细粒度音乐检索数据集,适合做音乐智能推荐的研究者

使用自然语言描述检索音乐已因对比音频-文本模型(如CLAP)取得进展,但现有系统仍局限于粗粒度语义查询。当描述涉及节拍、调性、和弦进行或节奏结构等细粒度音乐属性时,现有模型常无法准确检索。我们发现这一局限源于对比学习目标本身:尽管训练时使用长描述,但基于CLAP的模型仅有效利用前几个词元,丢弃了大量详细提示中的信息。为此,我们提出FIGMA(FIne-Grained Music RetrievAl),一种多视角对比架构,通过联合优化全局音频-文本对齐与帧级、词元级别的对齐,实现统一表示空间中高阶语义与细粒度音乐属性的共同捕捉。此外,我们正式定义细粒度音乐检索任务,并构建了包含38万组音乐-描述对的细粒度音乐描述数据集(FGMCaps),其测试集为1万条,均标注了节拍、调性、和弦进行、节拍数以及流派与情绪。大量实验证明,FIGMA在多个音乐检索基准上持续优于现有基于CLAP的模型,包括跨域评估,相对提升最高达73.3%。

原文摘要 · Abstract (English)

Retrieving music using natural language descriptions has improved with contrastive audio-text models such as CLAP, but current systems remain limited to coarse semantic queries. When descriptions specify fine-grained musical attributes such as tempo, key, chord progression, or rhythmic structure, existing models often fail to retrieve the correct audio. We show that this limitation stems from the contrastive learning objective itself: despite being trained on long captions, CLAP-based models effectively utilize only the first few tokens, discarding much of the information encoded in detailed prompts. Then, we propose FIGMA (FIne-Grained Music RetrievAl), a multi-view contrastive architecture that addresses this limitation by jointly optimizing global audio-text alignment and frame-level, token-wise alignment. This design enables FIGMA to capture both high-level semantic context and fine-grained musical attributes within a unified representation space. Moreover, we formalize the task of Fine-Grained Music Retrieval and construct Fine-Grained Music Caption dataset (FGMCaps), a large-scale dataset of 380K music-caption pairs for training along with a 10K test set, both annotated with tempo, key, chord progression, beat count, as well as genre and mood. Extensive experiments demonstrate that FIGMA consistently outperforms existing CLAP-based music retrieval models across multiple music retrieval benchmarks, including out-of-domain evaluations, with relative improvements of up to 73.3%.

音乐检索细粒度对比学习数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。