为蒙古语大模型评估构建分层基准,揭示理解短板。
MM-Eval: A Hierarchical Benchmark for Modern Mongolian Evaluation in LLMs
- 按语法、语义、知识、推理分层设计评测体系
- 所有模型语义任务表现差于语法,知识可部分迁移
- 适合低资源语言NLP研究者和多语言模型开发者
大型语言模型在高资源语言上表现优异,但在蒙古语等低资源语言上面临挑战。本文将能力分为语言能力(语法与语义)和认知能力(知识与推理),构建了基于《现代蒙古语教材Ⅰ》并融合WebQSP和MGSM数据集的专用评测集MM-Eval。在Qwen2-7B-Instruct、GLM4-9b-chat、Llama3.1-8B-Instruct、GPT-4和DeepseekV2.5等模型上的初步实验表明:1)所有模型在语法任务上表现优于语义任务,凸显深层语言理解的不足;2)知识类任务中等下降,表明模型能将高资源语言的知识部分迁移到低资源场景。MM-Eval包含569个语法、677个语义、344个知识和250个推理任务,为提升蒙古语等低资源语言的NLP与大模型研究提供重要参考。数据集已开源:https://github.com/joenahm/MM-Eval。
原文摘要 · Abstract (English)
Large language models (LLMs) excel in high-resource languages but face notable challenges in low-resource languages like Mongolian. This paper addresses these challenges by categorizing capabilities into language abilities (syntax and semantics) and cognitive abilities (knowledge and reasoning). To systematically evaluate these areas, we developed MM-Eval, a specialized dataset based on Modern Mongolian Language Textbook I and enriched with WebQSP and MGSM datasets. Preliminary experiments on models including Qwen2-7B-Instruct, GLM4-9b-chat, Llama3.1-8B-Instruct, GPT-4, and DeepseekV2.5 revealed that: 1) all models performed better on syntactic tasks than semantic tasks, highlighting a gap in deeper language understanding; and 2) knowledge tasks showed a moderate decline, suggesting that models can transfer general knowledge from high-resource to low-resource contexts. The release of MM-Eval, comprising 569 syntax, 677 semantics, 344 knowledge, and 250 reasoning tasks, offers valuable insights for advancing NLP and LLMs in low-resource languages like Mongolian. The dataset is available at https://github.com/joenahm/MM-Eval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。