arXiv:2601.13931cs.SDcs.IR2026-01中稿 · IEEE International…

提升音乐模型对否定语义的识别能力,让系统更懂‘无歌声’和‘有歌声’的区别。

Towards Effective Negation Modeling in Joint Audio-Text Models for Music

  • 从头训练CLAP模型,用带否定的文本增强数据和对比损失分离正负样本
  • 在百万首歌数据集上,否定识别准确率显著提升,检索性能基本不变
  • 适合需要精确理解音乐描述中否定词的研究者与应用开发者

联合音频-文本模型广泛用于音乐检索,但在处理否定等语义现象时表现不佳。否定是区分音乐元素存在与否(如“有歌声”与“无歌声”)的基础,但现有系统难以可靠建模。本文基于百万首歌数据集(Million Song Dataset)与LP-MusicCaps-MSD标注,从零开始训练CLAP模型,通过文本增强引入否定,并设计基于差异性的对比损失,使原始与否定描述在联合嵌入空间中明确分离。为评估进展,提出两种任务协议:将否定建模转化为检索与二分类任务。实验表明,两种方法单独或联合使用均显著改善否定处理能力,同时保持原有检索性能。关键指标提升在测试集上达12.7%(绝对),且不牺牲主流检索指标(mAP)。

原文摘要 · Abstract (English)

Joint audio-text models are widely used for music retrieval, yet they struggle with semantic phenomena such as negation. Negation is fundamental for distinguishing the absence (or presence) of musical elements (e.g., "with vocals" vs. "without vocals"), but current systems fail to represent this reliably. In this work, we investigate and mitigate this limitation by training CLAP models from scratch on the Million Song Dataset with LP-MusicCaps-MSD captions. We introduce negation through text augmentation and a dissimilarity-based contrastive loss, designed to explicitly separate original and negated captions in the joint embedding space. To evaluate progress, we propose two protocols that frame negation modeling as retrieval and binary classification tasks. Experiments demonstrate that both methods, individually and combined, improve negation handling while largely preserving retrieval performance.

音乐检索否定建模多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。