通过干预模型激活,无需微调即可提升多语言模型的跨语言对齐效果。
Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language Models
- 通过定位并操控特定神经元,引导模型生成更对齐的跨语言表示。
- 在跨语言检索任务中,top-1准确率最高提升2倍。
- 适合关注高效多语言建模与模型可解释性的研究者。
多语言大模型(mLLMs)中跨语言表征对齐是一种理想特性,有助于提升跨语言任务性能。传统对齐依赖微调,成本高且需大量语言数据,往往难以获取。一种数据高效的替代方案是模型干预——通过操纵模型激活来引导生成方向。本文分析了一种常见干预方法(寻找专家)对mLLMs跨语言表征对齐的影响。我们识别出针对特定语言需干预的神经元,并审视干预前后mLLMs的嵌入空间。结果表明,修改模型激活会改变其嵌入空间,增强跨语言对齐;进一步发现,这种空间变化可转化为下游检索任务性能提升,跨语言检索任务中top-1准确率最高提升2倍。
原文摘要 · Abstract (English)
Aligned representations across languages is a desired property in multilingual large language models (mLLMs), as alignment can improve performance in cross-lingual tasks. Typically alignment requires fine-tuning a model, which is computationally expensive, and sizable language data, which often may not be available. A data-efficient alternative to fine-tuning is model interventions -- a method for manipulating model activations to steer generation into the desired direction. We analyze the effect of a popular intervention (finding experts) on the alignment of cross-lingual representations in mLLMs. We identify the neurons to manipulate for a given language and introspect the embedding space of mLLMs pre- and post-manipulation. We show that modifying the mLLM's activations changes its embedding space such that cross-lingual alignment is enhanced. Further, we show that the changes to the embedding space translate into improved downstream performance on retrieval tasks, with up to 2x improvements in top-1 accuracy on cross-lingual retrieval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。