arXiv:2502.21228cs.CLcs.AI2025-02被引 28

构建多语言评测集,检验大模型跨语言知识迁移能力。

ECLeKTic: a Novel Challenge Set for Evaluation of Cross-Lingual Knowledge Transfer

  • 基于12种语言维基百科有无信息,设计跨语言知识题
  • 8个SOTA模型在跨语言迁移上表现不佳,准确率低
  • 适合评估大模型真实跨语言理解能力,非仅语言内预测

为实现多语言间公平性能,大语言模型需超越训练语言抽象知识。然而现有研究缺乏可靠方法评估其跨语言知识迁移能力。为此,我们提出ECLeKTic,一个面向多语言闭卷问答的评测集,以简单黑盒方式评估跨语言知识转移。具体地,利用12种语言维基百科文章的有无差异,识别出在某一语言中可能存在于预训练阶段但其他语言中缺失的信息。我们据此构建了覆盖所有语言的事实类问题集,要求模型必须实现跨语言知识迁移才能正确回答。我们评估了8个大语言模型,发现当前最先进模型即便在知识获取语言中能准确作答,仍难以有效将知识迁移到其他语言。

原文摘要 · Abstract (English)

To achieve equitable performance across languages, large language models (LLMs) must be able to abstract knowledge beyond the language in which it was learnt. However, the current literature lacks reliable ways to measure LLMs' capability of such cross-lingual knowledge transfer. To that end, we present ECLeKTic, a multilingual closed-book QA dataset that Evaluates Cross-Lingual Knowledge Transfer in a simple, black-box manner. Concretely, we used the presence and absence of Wikipedia articles in 12 languages to detect pieces of information that were likely available during pre-training in one of the languages but not in the others. We curate ECLeKTic as a set of fact-seeking questions over this kind of information, in all the different languages. Therefore, in order to solve ECLeKTic the model is required to transfer knowledge between languages. We evaluated 8 LLMs and showed that current SOTA models struggle to effectively share knowledge across languages, even if they can predict the answer for questions in the language in which the knowledge was acquired.

跨语言知识迁移评测集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。