arXiv:2510.12115cs.CL2025-10

研究双语医学模型如何跨语言学习知识,发现即使有高质量语料仍难迁移。

Tracing Multilingual Knowledge Acquisition Dynamics in Domain Adaptation: A Case Study of English-Japanese Biomedical Adaptation

  • 构建自适应评估方法AdaXEval,用同源双语语料生成测评数据
  • 130亿参数模型在跨语言知识迁移中仍表现不佳,尤其低资源时
  • 揭示知识从训练数据到模型的转化机制,适合多语言模型研究者

多语言领域适应(ML-DA)广泛用于将新领域的知识融入大语言模型(LLMs)。尽管已有诸多方法提升适应效果,但多语言知识获取的内在机制——即知识如何在单语内学习并跨语言传递——仍不清晰。这一空白导致低资源场景下性能不佳。本文研究了LLM在多语言领域适应中的学习动态。由于以往研究常使用知识覆盖不匹配的数据集进行训练与评估,本文提出AdaXEval:一种基于同一双语领域语料构建多项选择题数据集的自适应评估方法,可直接观测多语言知识获取过程。通过持续训练不同数据策略的模型,追踪其获取领域事实的过程,并定位从训练数据到知识的转化机制。在130亿参数的英日双语模型上的实验表明,即使拥有高质量双语语料,跨语言知识迁移仍面临挑战。代码已公开。

原文摘要 · Abstract (English)

Multilingual domain adaptation (ML-DA) is widely used to learn new domain knowledge across languages into large language models (LLMs). Although many methods have been proposed to improve domain adaptation, the mechanisms of multilingual knowledge acquisition, how domain knowledge is learned within a language and transferred across languages, remain underexplored. This gap leads to suboptimal performance, particularly in low-resource settings. This work examines the learning dynamics of LLMs during ML-DA. Because prior ML-DA studies often train and evaluate on datasets with mismatched knowledge coverage, we propose AdaXEval, an adaptive evaluation method that builds multiple-choice QA datasets from the same bilingual domain corpus used for training, thereby directly studying multilingual knowledge acquisition. Through continual training of LLMs with diverse data recipes, we track how LLMs acquire domain facts and pinpoint the mechanism behind the transformation process from domain training data to knowledge. Our experiments on a 13B English-Japanese bilingual LLM reveal that cross-lingual transfer remains challenging despite a high-quality bilingual corpus. The code has been released.

多语言知识迁移大模型领域适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。