arXiv:2602.14127cs.SD2026-02

MUKA通过多核融合提升音频语言模型少样本适应能力

MUKA: Multi Kernel Audio Adaptation Of Audio-Language Models

  • 用多核方法融合细粒度与全局语义表示
  • 11个数据集上超越现有无训练适配方法
  • 无需额外训练,适合快速部署场景

多模态基础模型展现出强大泛化能力,但在少样本设置下高效适配仍具挑战。本文研究基于训练和免训练的大型音频-语言模型(ALMs)少样本适配。提出MUKA框架,结合指令微调模型(如Pengi)的细粒度上下文表示与对比预训练方法(如CLAP)的全局语义表示,通过构建乘积核对齐局部相似性与全局语义,增强表征能力并保持核方法理论保证,且无需额外训练。在11个多样化音频数据集上的大量实验表明,MUKA在免训练方法中达到最先进性能,甚至在多个场景超越训练型适配器,实现了适应性与效率的出色平衡。

原文摘要 · Abstract (English)

Multimodal foundation models have demonstrated impressive generalization capabilities, yet efficiently adapting them to new tasks in a few-shot setting remains a critical challenge. In this work, we investigate the few-shot adaptation of Large Audio-Language Models (ALMs) through both training-based and training-free approaches. We introduce MUKA, a multi-kernel adaptation framework that combines the fine-grained, context-dependent representations of instruction-tuning based models like Pengi with the global semantic representations of contrastive pretraining methods like CLAP. By constructing a product kernel that aligns local similarity with global semantics, MUKA enhances representational power while preserving the theoretical guarantees of kernel methods and avoiding additional training. Extensive experiments across 11 diverse audio datasets demonstrate that MUKA achieves state-of-the-art performance among training-free methods and even surpasses training-based adapters in several scenarios, offering a compelling balance between adaptability and efficiency.

音频语言模型少样本学习多核方法免训练适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。