Lens让大模型更懂多语言,不丢英语能力还省算力。
Lens: Rethinking Multilingual Enhancement for Large Language Models
- 利用语言内表示空间,分两步提升多语能力
- 三款英文主导模型多语性能显著提升,英语能力不变
- 比现有方法更省资源,适合需要多语支持的场景
随着全球对多语言大语言模型(LLMs)需求增长,大多数模型仍过度聚焦英语,限制了非英语使用者获取先进AI的能力。当前增强多语能力的方法主要依赖数据驱动的后训练技术,如多语言指令微调或持续预训练,但存在资源消耗高、副作用加剧及核心语言能力灾难性遗忘等问题。为此,我们提出Lens,一种新方法,通过利用大模型内部的语言表征空间来增强多语能力。Lens在两个子空间中运作:语言无关子空间中,将目标语言与中心语言对齐,继承强语义表征;语言特定子空间中,分离目标语言与中心语言,保留语言特异性。在三个以英语为中心的LLM上实验表明,Lens显著提升了多语性能,同时保持了模型的英语能力,相较现有后训练方法实现更好效果且计算成本更低。
原文摘要 · Abstract (English)
As global demand for multilingual large language models (LLMs) grows, most LLMs still remain overly focused on English, leading to the limited access to advanced AI for non-English speakers. Current methods to enhance multilingual capabilities largely rely on data-driven post-training techniques, such as multilingual instruction tuning or continual pre-training. However, these approaches exhibit significant limitations, including high resource cost, exacerbation of off-target issue and catastrophic forgetting of central language abilities. To this end, we propose Lens, a novel approach that enhances multilingual capabilities by leveraging LLMs' internal language representation spaces. Lens operates on two subspaces: the language-agnostic subspace, where it aligns target languages with the central language to inherit strong semantic representations, and the language-specific subspace, where it separates target and central languages to preserve linguistic specificity. Experiments on three English-centric LLMs show that Lens significantly improves multilingual performance while maintaining the model's English proficiency, achieving better results with less computational cost compared to existing post-training approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。