LLM在多语言表现中受语系结构影响,但训练数据不平衡是主导因素。
Are the LLMs Capable of Maintaining at Least the Language Genus?
- 通过跨语系对比,检验模型是否偏好相关语言
- 语系内知识一致性高于语系间,但受训练资源制约
- 不同模型家族有差异化的多语言策略,适合研究语系影响者关注
大型语言模型(LLMs)在多语言行为上表现出显著差异,但语系结构对其影响尚未充分探索。本文基于MultiQ数据集扩展分析,考察模型在提示语言不一致时是否倾向切换至语系相关的语言,并检验语系内与语系间的知识一致性。结果表明,语系层面效应存在,但强烈依赖训练资源的可获得性。不同模型家族展现出不同的多语言策略。研究发现,LLMs能编码部分语系结构特征,但训练数据分布不均仍是决定其多语言表现的主要因素。
原文摘要 · Abstract (English)
Large Language Models (LLMs) display notable variation in multilingual behavior, yet the role of genealogical language structure in shaping this variation remains underexplored. In this paper, we investigate whether LLMs exhibit sensitivity to linguistic genera by extending prior analyses on the MultiQ dataset. We first check if models prefer to switch to genealogically related languages when prompt language fidelity is not maintained. Next, we investigate whether knowledge consistency is better preserved within than across genera. We show that genus-level effects are present but strongly conditioned by training resource availability. We further observe distinct multilingual strategies across LLMs families. Our findings suggest that LLMs encode aspects of genus-level structure, but training data imbalances remain the primary factor shaping their multilingual performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。