揭示大模型跨语言理解中表示对齐的关键作用
Can you map it to English? The Role of Cross-Lingual Alignment in Multilingual Performance of LLMs
- 提出实例级对齐度量方法,分析24种语言的跨语言表示对齐
- 发现非英语输入在中间层与英文对齐度低时易出错
- 通过激活修补验证对齐是跨语言正确性的因果因素
大语言模型虽主要在英语数据上训练,却能回答多种语言的问题,但其泛化机制仍不清晰。本文研究模型将非英语输入表示对齐到英语的能力如何影响自然语言理解(NLU)任务表现。我们引入判别性对齐指数($\ ext{DALI}$),量化24种非英语语言在三个不同NLU任务上的实例级对齐程度。结果表明,错误的NLU预测与模型中间层中非英语表示与英语对齐度较低显著相关。通过激活修补实验,我们发现用对应英语激活替换非英语输入的中间层激活,可纠正非英语任务中的错误,证明了表示(误)对齐在跨语言正确性中的因果作用。
原文摘要 · Abstract (English)
Large language models (LLMs) can answer prompts in many languages, despite being trained predominantly on English; yet, the mechanisms driving this generalization remain poorly understood. This work asks: How does an LLM's ability to align representations of non-English inputs to English impact its performance on natural language understanding (NLU) tasks? We study the role of representation alignment in instance-level task decisions, complementing prior analyses conducted both at the language level and task-independently. We introduce the Discriminative Alignment Index ($\DALI$) to quantify instance-level alignment across 24 languages other than English and three distinct NLU tasks. Results show that incorrect NLU predictions are strongly associated with lower representation alignment with English in the model's middle layers. Through activation patching, we show that incorrect predictions in languages other than English can be fixed by patching their parallel English activations in the middle layers, thereby demonstrating the causal role of representation (mis)alignment in cross-lingual correctness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。