arXiv:2410.02271cs.SDcs.AI2024-10被引 10

让模型理解5分钟音乐与长文本的对应关系,提升跨模态检索准确率。

CoLLAP: Contrastive Long-form Language-Audio Pretraining with Musical Temporal Structure Augmentation

  • 用音乐时间结构增强长音频描述,构建51.3K对长音频-文本数据
  • 通过分段嵌入与注意力机制,实现跨模态时间对齐,检索准确率显著提升
  • 适用于音乐信息检索、跨模态生成等需要长上下文理解的任务

建模音频波形的时间特性对表示学习至关重要。我们提出对比式长文本语言-音频预训练(CoLLAP),将输入音频感知窗口扩展至5分钟,语言描述超过250词,同时支持跨模态与时间动态的对比学习。利用近期音乐大模型为完整歌曲生成长文本描述,并融合音乐时间结构,从大规模 AudioSet 训练集收集了51.3K组音频-文本对,平均音频长度达288秒。提出一种新颖的对比学习架构,将语言表示与结构化音频表示融合:将每首歌分割为片段并提取嵌入,通过注意力机制捕捉多模态时间相关性,自动加权并增强最终融合得分,以改善对比对齐。最后,开发了两种基于不同骨干语言模型的CoLLAP变体。在多个长文本音乐-文本检索数据集上,实验表明其检索准确率持续优于基线模型。同时证明预训练的CoLLAP模型可迁移至多种音乐信息检索任务,适用于异构长上下文多模态场景。

原文摘要 · Abstract (English)

Modeling temporal characteristics plays a significant role in the representation learning of audio waveform. We propose Contrastive Long-form Language-Audio Pretraining (\textbf{CoLLAP}) to significantly extend the perception window for both the input audio (up to 5 minutes) and the language descriptions (exceeding 250 words), while enabling contrastive learning across modalities and temporal dynamics. Leveraging recent Music-LLMs to generate long-form music captions for full-length songs, augmented with musical temporal structures, we collect 51.3K audio-text pairs derived from the large-scale AudioSet training dataset, where the average audio length reaches 288 seconds. We propose a novel contrastive learning architecture that fuses language representations with structured audio representations by segmenting each song into clips and extracting their embeddings. With an attention mechanism, we capture multimodal temporal correlations, allowing the model to automatically weigh and enhance the final fusion score for improved contrastive alignment. Finally, we develop two variants of the CoLLAP model with different types of backbone language models. Through comprehensive experiments on multiple long-form music-text retrieval datasets, we demonstrate consistent performance improvement in retrieval accuracy compared with baselines. We also show the pretrained CoLLAP models can be transferred to various music information retrieval tasks, with heterogeneous long-form multimodal contexts.

多模态预训练音乐理解长序列建模对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。