评测多语言大模型对波罗的海与北欧历史的知识掌握情况
Towards Multilingual LLM Evaluation for Baltic and Nordic languages: A study on Lithuanian History
- 用多语种历史问答数据集测试多个大模型在立陶宛和泛历史知识上的表现
- GPT-4o表现最佳,开源模型中QWEN2.5 72b和LLaMa3.1 70b表现较好但对立陶宛语支持弱
- 专精北欧语系的微调模型未超越通用多语言模型,文化关联不必然提升性能
本文在多选题问答任务中评估了多语言大语言模型(LLMs)对立陶宛及普遍历史知识的掌握程度。模型在包含立陶宛、波罗的海、北欧及其他语言(英语、乌克兰语、阿拉伯语)的历史问题数据集上进行测试,以评估具有文化和历史关联性的语言群体间知识共享能力。测试模型包括GPT-4o、LLaMa3.1 8b与70b、QWEN2.5 7b与72b、Mistral Nemo 12b、LLaMa3 8b、Mistral 7b、LLaMa3.2 3b,以及针对北欧语言微调的GPT-SW3和LLaMa3 8b。结果表明,GPT-4o在所有语言组中均表现最优,且在波罗的海与北欧语言上略优。较大的开源模型如QWEN2.5 72b和LLaMa3.1 70b表现良好,但在立陶宛语相关任务上对齐度较弱。较小模型(如Mistral Nemo 12b、LLaMa3.2 3b、QWEN 7B、LLaMa3.1 8B、LLaMa3 8b)在立陶宛相关任务上存在差距,但在北欧及其他语言上表现更好。北欧微调模型未超越通用多语言模型,表明仅靠文化或历史关联性不足以提升性能。
原文摘要 · Abstract (English)
In this work, we evaluated Lithuanian and general history knowledge of multilingual Large Language Models (LLMs) on a multiple-choice question-answering task. The models were tested on a dataset of Lithuanian national and general history questions translated into Baltic, Nordic, and other languages (English, Ukrainian, Arabic) to assess the knowledge sharing from culturally and historically connected groups. We evaluated GPT-4o, LLaMa3.1 8b and 70b, QWEN2.5 7b and 72b, Mistral Nemo 12b, LLaMa3 8b, Mistral 7b, LLaMa3.2 3b, and Nordic fine-tuned models (GPT-SW3 and LLaMa3 8b). Our results show that GPT-4o consistently outperformed all other models across language groups, with slightly better results for Baltic and Nordic languages. Larger open-source models like QWEN2.5 72b and LLaMa3.1 70b performed well but showed weaker alignment with Baltic languages. Smaller models (Mistral Nemo 12b, LLaMa3.2 3b, QWEN 7B, LLaMa3.1 8B, and LLaMa3 8b) demonstrated gaps with LT-related alignment with Baltic languages while performing better on Nordic and other languages. The Nordic fine-tuned models did not surpass multilingual models, indicating that shared cultural or historical context alone does not guarantee better performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。