arXiv:2505.16385cs.CL2025-05

发现语言模型跨语言能力的关键机制,通过语义枢纽提升翻译效果。

Semantic Pivots Enable Cross-Lingual Transfer in Large Language Models

  • 通过词级翻译任务分析模型中间层,发现共现与语义枢纽两种行为
  • 基于预训练数据中的语义枢纽重构数据集,显著提升跨语言性能
  • 为理解模型跨语言机制提供新视角,适合语言模型优化研究者

大型语言模型在跨语言任务中表现优异,但其能力来源尚不明确。为准确量化该能力,我们提出词级跨语言翻译任务,并追踪模型在该任务中各中间层的输出。我们识别并区分了模型前向传播中的两种行为:共现行为与语义枢纽行为。研究表明,这两种行为分别源于词汇共现频率及预训练数据中的语义枢纽。为应用该发现,我们构建了一个以高比例语义枢纽为核心的重构预训练数据集。实验验证了该方法在提升跨语言能力上的有效性。本研究深化了对大模型跨语言能力可解释性的理解,并提出了可复用的改进策略。

原文摘要 · Abstract (English)

Large language models (LLMs) demonstrate remarkable ability in cross-lingual tasks. Understanding how LLMs acquire this ability is crucial for their interpretability. To quantify the cross-lingual ability of LLMs accurately, we propose a Word-Level Cross-Lingual Translation Task. To find how LLMs learn cross-lingual ability, we trace the outputs of LLMs' intermediate layers in the word translation task. We identify and distinguish two distinct behaviors in the forward pass of LLMs: co-occurrence behavior and semantic pivot behavior. We attribute LLMs' two distinct behaviors to the co-occurrence frequency of words and find the semantic pivot from the pre-training dataset. Finally, to apply our findings to improve the cross-lingual ability of LLMs, we reconstruct a semantic pivot-aware pre-training dataset using documents with a high proportion of semantic pivots. Our experiments validate the effectiveness of our approach in enhancing cross-lingual ability. Our research contributes insights into the interpretability of LLMs and offers a method for improving LLMs' cross-lingual ability.

语言模型跨语言可解释性语义枢纽

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。