通过学习定向向量提升语音大模型在跨域场景下的表现
SALSA: Speech Aware LLM Adaptation via Learned Steering Activation Vectors
- 用监督目标直接优化分层定向向量,不依赖对比激活差
- 在儿童语音等场景下相对零样本提升最高达46.8%
- 调整编码器后层比调整语言模型主干更有效
语音感知的大语言模型在域外设置下泛化能力通常较差。我们提出SALSA(基于学习定向激活的语音感知大模型适配),一种轻量级适配方法,通过学习分层定向向量实现改进。不同于依赖对比激活差异的常见方法,SALSA采用监督目标直接优化定向向量。在儿童语音、多语言语音及中英文混杂语音等多个基准上,SALSA显著优于零样本推理和语音上下文学习基线,相对零样本最高提升46.8%。分析表明,对编码器尤其是后层进行定向调节,比对大模型主干调节更有效。这说明定向调节通过使高层声学与音位表征更好地对齐预训练语言模型表示空间,从而提升下游语音识别性能,而非修改解码器本身。
原文摘要 · Abstract (English)
Speech-aware large language models often generalize poorly to out-of-domain settings. We propose SALSA (Speech-Aware LLM Adaptation via Learned Steering Activations), a lightweight adaptation method that learns layer-wise steering vectors. Unlike commonly used steering approaches that rely on contrastive activation differences, SALSA directly optimizes steering vectors using a supervised objective. Across children's speech, multilingual speech, and Mandarin-English code-switching benchmarks, SALSA substantially improves performance over zero-shot inference and speech in-context learning baselines, achieving up to 46.8% relative improvements over zero-shot. Analysis further demonstrates that steering the encoder, particularly the later layers, is more effective than steering the LLM backbone. These findings suggest that steering improves downstream ASR performance by adapting higher-level acoustic and phonetic representations to better align with the pretrained language model representation space, rather than by modifying the decoder itself.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。