arXiv:2601.20867cs.SDcs.AI2026-01ACL被引 1

通过语义扩展增强音频-语言模型提示的泛化能力

Generalizable Prompt Tuning for Audio-Language Models via Semantic Expansion

  • 用大模型生成语义邻近词,构建更稳定的提示嵌入空间
  • 在多个基线方法上提升跨数据集泛化性能,且推理开销不变
  • 适合需要高泛化性的音频-语言任务研究者

提示调优在视觉-语言模型中已取得显著进展,并被逐步引入音频-语言模型(ALMs)。然而,其在ALMs中的泛化能力仍缺乏系统研究。我们发现,传统提示调优在ALMs中同样存在基础-新类别权衡问题,根源在于嵌入空间语义结构被破坏。为此,提出语义扩展提示调优(SEPT)——一种即插即用框架,通过引入大语言模型生成的语义邻居,显式正则化提示嵌入空间。SEPT设计了一种带边界约束的语义扩展损失,促进类内紧凑性与类间可分性,从而增强提示嵌入空间的语义结构。为全面评估,建立了首个针对ALMs提示泛化的基准设置,涵盖基类到新类泛化及跨数据集迁移能力。大量实验表明,SEPT在多个提示调优基线上持续提升泛化性能,且推理时计算成本保持不变。

原文摘要 · Abstract (English)

Prompt tuning has achieved remarkable progress in vision-language models (VLMs) and is recently being adopted for audio-language models (ALMs). However, its generalization ability in ALMs remains largely underexplored. We observe that conventional prompt tuning for ALMs also suffers from the Base-New Tradeoff, and we identify that this issue stems from the disrupted semantic structure of the embedding space. To address this issue, we propose Semantically Expanded Prompt Tuning (SEPT)-a plug-and-play framework that explicitly regularizes the prompt embedding space by incorporating semantic neighbors generated by large language models. SEPT introduces a novel semantic expansion loss with margin constraints that promote intra-class compactness and inter-class separability, thereby enhancing the semantic structure of the prompt embedding space. For comprehensive evaluation, we establish the first benchmark setup for prompt generalization in ALMs, covering both base-to-new generalization and cross-dataset transferability. Extensive experiments demonstrate that SEPT consistently improves generalization performance across multiple prompt tuning baselines, while maintaining computational cost during inference.

提示调优音频-语言泛化能力语义结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。