大模型跨语言概念空间隐含对齐,影响多语泛化效果。
Concept Space Alignment in Multilingual LLMs
- 发现大模型中不同语言概念向量存在高质量线性对齐。
- 模型越大,跨语言对齐越强,但抽象概念和相似语系效果更佳。
- 提示词嵌入对齐不如词嵌入线性,可能破坏隐式对齐机制。
多语言大语言模型在跨语言上表现出一定泛化能力。我们推测这源于隐式的向量空间对齐。评估发现,更大模型在不同语言对应概念间展现出极高的线性对齐质量。实验表明,多语言大模型存在两个常见弱点:跨语言泛化对语言类型相似的语言更有效,且对抽象概念表现更好。对于部分模型(如 Llama-2 系列),基于提示的嵌入比词嵌入对齐效果更优,但其投影关系更非线性——这一现象几乎在所有模型族中成立,说明提示方法可能部分破坏了隐式学习到的对齐结构。
原文摘要 · Abstract (English)
Multilingual large language models (LLMs) seem to generalize somewhat across languages. We hypothesize this is a result of implicit vector space alignment. Evaluating such alignment, we see that larger models exhibit very high-quality linear alignments between corresponding concepts in different languages. Our experiments show that multilingual LLMs suffer from two familiar weaknesses: generalization works best for languages with similar typology, and for abstract concepts. For some models, e.g., the Llama-2 family of models, prompt-based embeddings align better than word embeddings, but the projections are less linear -- an observation that holds across almost all model families, indicating that some of the implicitly learned alignments are broken somewhat by prompt-based methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。