研究35个语言模型,发现本地化LLM需针对性训练特定能力。
Why We Build Local Large Language Models: An Observational Analysis from 35 Japanese and Multilingual LLMs
- 通过评测和主成分分析,提取日语模型的能力因子
- 英语训练可提升日语学术任务表现,但非所有能力都需日语数据
- 日语问答与翻译能力依赖日语数据训练,具本地化特征
为何要构建本地大语言模型?本地模型应从目标语言中学习什么?其他语言的能力能否迁移?是否存在语言特异性扩展规律?为探究这些问题,我们对35个日语、英语及多语言LLM在19个日英评估基准上进行了评测,以日语为本地语言。采用观察性方法,分析基准得分相关性,并对得分进行主成分分析(PCA),提取本地LLM的“能力因子”。结果发现,在英语文本上训练可提升日语学术任务(JMMLU)表现;而日语代码生成、算术推理、常识判断和阅读理解等任务,无需专门日语训练即可获得良好性能。相反,日语知识问答和英日翻译能力需日语文本训练才能提升,表明这两项能力可视为日语专属能力。此外,验证了日语能力随日语文本计算预算增加而增长,存在语言特异性扩展规律。
原文摘要 · Abstract (English)
Why do we build local large language models (LLMs)? What should a local LLM learn from the target language? Which abilities can be transferred from other languages? Do language-specific scaling laws exist? To explore these research questions, we evaluated 35 Japanese, English, and multilingual LLMs on 19 evaluation benchmarks for Japanese and English, taking Japanese as a local language. Adopting an observational approach, we analyzed correlations of benchmark scores, and conducted principal component analysis (PCA) on the scores to derive \textit{ability factors} of local LLMs. We found that training on English text can improve the scores of academic subjects in Japanese (JMMLU). In addition, it is unnecessary to specifically train on Japanese text to enhance abilities for solving Japanese code generation, arithmetic reasoning, commonsense, and reading comprehension tasks. In contrast, training on Japanese text could improve question-answering tasks about Japanese knowledge and English-Japanese translation, which indicates that abilities for solving these two tasks can be regarded as \textit{Japanese abilities} for LLMs. Furthermore, we confirmed that the Japanese abilities scale with the computational budget for Japanese text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。