现有音频语言模型对文本变化不敏感,新方法提升检索鲁棒性
Do Audio-Language Models Understand Linguistic Variations?
- 用多视角对比学习,把同场景的改写句视为等价视图
- 在多个数据集上使检索准确率提升0.8%至13%
- 适合需要稳定文本查询响应的音频检索应用
开放词汇音频语言模型(ALMs),如对比语言音频预训练(CLAP),利用自然语言查询进行音视频检索,是新兴范式。本文首次在多个基准上开展受控实验,表明现有ALMs难以泛化到文本查询中的语言变化。为此,我们提出RobustCLAP,一种新型且计算高效的方案,使音频-语言表征对语言变化具有鲁棒性。具体地,通过引入多视角对比学习目标,将改写句视为同一音频场景的不同视角,并以此训练模型。该方法在多个基准上将CLAP的文本到音频检索性能提升0.8%至13%,显著增强对语言变化的鲁棒性。
原文摘要 · Abstract (English)
Open-vocabulary audio language models (ALMs), like Contrastive Language Audio Pretraining (CLAP), represent a promising new paradigm for audio-text retrieval using natural language queries. In this paper, for the first time, we perform controlled experiments on various benchmarks to show that existing ALMs struggle to generalize to linguistic variations in textual queries. To address this issue, we propose RobustCLAP, a novel and compute-efficient technique to learn audio-language representations agnostic to linguistic variations. Specifically, we reformulate the contrastive loss used in CLAP architectures by introducing a multi-view contrastive learning objective, where paraphrases are treated as different views of the same audio scene and use this for training. Our proposed approach improves the text-to-audio retrieval performance of CLAP by 0.8%-13% across benchmarks and enhances robustness to linguistic variation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。