通过语义扩展增强音频-语言模型提示的泛化能力
Generalizable Prompt Tuning for Audio-Language Models via Semantic Expansion
- 用大模型生成语义邻近词,构建更稳定的提示嵌入空间
- 在多个基线方法上提升跨数据集泛化性能,且推理开销不变
- 适合需要高泛化性的音频-语言任务研究者
提示调优在视觉-语言模型中已取得显著进展,并被逐步引入音频-语言模型(ALMs)。然而,其在ALMs中的泛化能力仍缺乏系统研究。我们发现,传统提示调优在ALMs中同样存在基础-新类别权衡问题,根源在于嵌入空间语义结构被破坏。为此,提出语义扩展提示调优(SEPT)——一种即插即用框架,通过引入大语言模型生成的语义邻居,显式正则化提示嵌入空间。SEPT设计了一种带边界约束的语义扩展损失,促进类内紧凑性与类间可分性,从而增强提示嵌入空间的语义结构。为全面评估,建立了首个针对ALMs提示泛化的基准设置,涵盖基类到新类泛化及跨数据集迁移能力。大量实验表明,SEPT在多个提示调优基线上持续提升泛化性能,且推理时计算成本保持不变。
原文摘要 · Abstract (English)
Prompt tuning has achieved remarkable progress in vision-language models (VLMs) and is recently being adopted for audio-language models (ALMs). However, its generalization ability in ALMs remains largely underexplored. We observe that conventional prompt tuning for ALMs also suffers from the Base-New Tradeoff, and we identify that this issue stems from the disrupted semantic structure of the embedding space. To address this issue, we propose Semantically Expanded Prompt Tuning (SEPT)-a plug-and-play framework that explicitly regularizes the prompt embedding space by incorporating semantic neighbors generated by large language models. SEPT introduces a novel semantic expansion loss with margin constraints that promote intra-class compactness and inter-class separability, thereby enhancing the semantic structure of the prompt embedding space. For comprehensive evaluation, we establish the first benchmark setup for prompt generalization in ALMs, covering both base-to-new generalization and cross-dataset transferability. Extensive experiments demonstrate that SEPT consistently improves generalization performance across multiple prompt tuning baselines, while maintaining computational cost during inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。