用跨语言对齐分数预测多语分类与翻译性能,发现英语仍是关键枢纽。
Predicting Multilingual Classification and Translation Performance of LLMs with Cross-Lingual Alignment -- Is English Enough?

- 对比27种跨语言对齐评分方法,评估其预测能力。
- 英语对齐度可媲美甚至优于源-目标对齐,预示模型以英语为内嵌枢纽。
- 提出基于PMI的翻译指标,适用于跨语言性能分析,关联性强。
多语言大语言模型在非英语分类任务上的表现,与其目标语言表示与英语的对齐程度正相关。已有多种跨语言对齐(CLA)评分方法被提出,结合不同嵌入提取方式。本文系统比较了27种CLA评分变体,分析其差异及在三大下游任务中的预测效果。值得注意的是,尽管大语言模型广泛用于生成任务如机器翻译,但以往研究几乎仅聚焦于分类任务。因此,本文进一步探究CLA评分是否同样能预测翻译性能。为此,我们提出一种基于PMI的翻译评价指标,降低对目标语言的依赖,并与chrF具有强相关性。结果表明,以英语为基准的CLA评分,对翻译质量的预测能力可媲美甚至超越源-目标间的CLA,为模型内部使用英语作为枢纽语言提供了新证据。
原文摘要 · Abstract (English)
Multilingual large language models (LLMs) have been shown to perform better on non-English classification tasks when the representations of the given language are more aligned to English within the model. Several cross-lingual alignment (CLA) scores have been proposed for use with LLMs, along with multiple approaches for extracting embeddings from the models. We provide a comparative analysis of 27 CLA score variants, examining how they differ and how well each predicts downstream performance across three tasks. Crucially, while LLMs are widely used for generative tasks such as machine translation, prior work has focused almost exclusively on classification. We therefore investigate whether CLA scores are similarly predictive of translation performance. To enable computing correlations across target languages, we propose a PMI-based translation metric, which is less dependent on the target language and correlates strongly with chrF. We find that CLA with English predicts translation quality comparably to or better than source-target CLA, providing new evidence that LLMs use English as an internal pivot language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。