arXiv:2511.14774cs.CLcs.AI2025-11ACL被引 1

构建动态基准,精准评估大模型跨语言知识迁移能力

LiveCLKTBench: Towards Reliable Evaluation of Cross-Lingual Knowledge Transfer in Multilingual LLMs

  • 通过真实时间敏感实体生成多语言事实题,隔离真正迁移知识
  • 发现跨语言迁移受语言距离影响显著,且方向不对称
  • 适合研究多语言模型性能与可解释性的学者使用

评估大语言模型中的跨语言知识迁移极具挑战性,因为目标语言的正确答案可能源于真正的知识迁移,也可能来自预训练期间的先前接触。我们提出 LiveCLKTBench,一个专门设计用于隔离并测量跨语言知识迁移的自动化生成管道。该管道从现实世界领域中识别出自包含、具有时间敏感性的知识实体,基于时间发生性进行筛选,并与模型知识进行验证。将这些有效实体的文档用于生成事实性问题,并翻译成多种语言,以评估跨语言边界的知识迁移能力。利用 LiveCLKTBench,我们在五种语言上评估了多个 LLM,发现跨语言迁移受语言距离强烈影响,且常呈现方向不对称性。虽然更大的模型能提升迁移效果,但收益随规模增长而递减,并在不同领域间存在差异。这些发现为多语言迁移提供了新见解,并证明了 LiveCLKTBench 作为未来研究可靠基准的价值。

原文摘要 · Abstract (English)

Evaluating cross-lingual knowledge transfer in large language models is challenging, as correct answers in a target language may arise either from genuine transfer or from prior exposure during pre-training. We present LiveCLKTBench, an automated generation pipeline specifically designed to isolate and measure cross-lingual knowledge transfer. Our pipeline identifies self-contained, time-sensitive knowledge entities from real-world domains, filters them based on temporal occurrence, and verifies them against the model's knowledge. The documents of these valid entities are then used to generate factual questions, which are translated into multiple languages to evaluate transferability across linguistic boundaries. Using LiveCLKTBench, we evaluate several LLMs across five languages and observe that cross-lingual transfer is strongly influenced by linguistic distance and often asymmetric across language directions. While larger models improve transfer, the gains diminish with scale and vary across domains. These findings provide new insights into multilingual transfer and demonstrate the value of LiveCLKTBench as a reliable benchmark for future research.

多语言模型知识迁移评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。