用文本提示增强音频表征,让模型更懂复杂音频语义。
Text-Prompted CLAP: Learning Text-Conditioned Audio Representations via Contrastive Learning

- 通过交叉注意力融合文本提示与音频特征,建模跨模态语义。
- 在音频问答任务上表现媲美大模型,且提升基础模型检索性能。
- 参数高效,适合音乐等属性导向的音频检索应用。
对比语言-音频预训练(CLAP)在共享嵌入空间中学习对齐的文本与音频表征。然而,各模态独立编码限制了其在复杂音频理解与检索任务中的跨模态语义建模能力。为此,本文提出文本提示增强的CLAP(TP-CLAP),一种参数高效的CLAP扩展,引入基于交叉注意力的融合模块,将文本提示融入音频特征。TP-CLAP采用音频多选题问答(AMCQA)框架训练,通过对比学习使文本条件下的音频表征与正确选项的文本嵌入对齐。实验表明,TP-CLAP在音频问答任务上表现媲美显著更大的音频大模型,同时提升基础CLAP在传统音频-文本检索和零样本分类基准上的性能。进一步微调后,该模型在属性聚焦的音频到音频检索任务中持续优于标准CLAP基线,尤其在音乐检索中表现突出。
原文摘要 · Abstract (English)
Contrastive Language-Audio Pretraining (CLAP) learns aligned text and audio representations in a shared embedding space. However, independent encoding of each modality limits its ability to model cross-modal semantics in complex audio understanding and retrieval tasks. To address this limitation, this paper proposes Text-Prompted CLAP (TP-CLAP), a parameter-efficient extension of CLAP that introduces a cross-attention-based fusion module to incorporate textual prompts into audio features. TP-CLAP is trained using an audio multiple-choice question answering (AMCQA) framework, where it learns to align text-conditioned audio representations with text embeddings of correct answer choices via contrastive learning. Experiments demonstrate that TP-CLAP performs competitively with substantially larger audio-LLMs on audio question answering, while also improving the base CLAP model on conventional audio-text retrieval and zero-shot classification benchmarks. The learned representations are further fine-tuned for attribute-focused audio-to-audio retrieval, showing that TP-CLAP consistently outperforms the standard CLAP baseline in music retrieval tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。