arXiv:2412.17306cs.SDcs.CV2024-12中稿 · ICASSP 2025被引 3

通过多一致性引导提升无标注音频的零样本音频-语言模型性能

Multiple Consistency-guided Test-Time Adaptation for Contrastive Audio-Language Models with Unlabeled Audio

  • 利用上下文与领域词元的一致性引导优化提示学习
  • 在12个下游任务上平均提升4.41%(最高7.50%)零样本性能
  • 适合需要无标签测试时适应的音频-语言模型研究者

预训练音频-语言模型(ALMs)具备出色的零样本泛化能力,测试时自适应(TTA)方法旨在不依赖标注的情况下提升域内表现。然而,现有针对零样本分类的ALMs TTA方法常陷入错误预测。为此,本文提出一种无需标注的多一致性引导提示学习方法:首先,在模型的上下文词元与领域词元上引入一致性约束;其次,对单个测试样本的多个增强视图间及不同测试样本间实施一致性与对比学习;最后,构建了对应的端到端学习框架。在跨12个领域的12个下游任务上评估,所提方法相比最先进模型平均提升4.41%(最高达7.50%)的零样本性能。

原文摘要 · Abstract (English)

One fascinating aspect of pre-trained Audio-Language Models (ALMs) learning is their impressive zero-shot generalization capability and test-time adaptation (TTA) methods aiming to improve domain performance without annotations. However, previous test time adaptation (TTA) methods for ALMs in zero-shot classification tend to be stuck in incorrect model predictions. In order to further boost the performance, we propose multiple guidance on prompt learning without annotated labels. First, guidance of consistency on both context tokens and domain tokens of ALMs is set. Second, guidance of both consistency across multiple augmented views of each single test sample and contrastive learning across different test samples is set. Third, we propose a corresponding end-end learning framework for the proposed test-time adaptation method without annotated labels. We extensively evaluate our approach on 12 downstream tasks across domains, our proposed adaptation method leads to 4.41% (max 7.50%) average zero-shot performance improvement in comparison with the state-of-the-art models.

音频-语言模型测试时适应无监督学习对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。