构建中文少数民族语言大模型评测基准,填补低资源语言评估空白
MiLiC-Eval: Benchmarking Multilingual LLMs for China's Minority Languages
- 设计涵盖9类任务、2.4万条数据的多语言评测集
- 开源模型在多文字系统任务中表现差,语法复杂任务尤为薄弱
- 为少数民族语言研究提供细粒度评估工具,适合语言技术开发者
大型语言模型在高资源语言上表现优异,但在中国少数民族语言(如藏语、维吾尔语、哈萨克语、蒙古语)等低资源语言上表现不佳。为系统追踪这些语言的发展进展,我们提出MiLiC-Eval,一个面向中国少数民族语言的评测基准,包含24,000个实例,覆盖9个任务。该基准聚焦于代表性不足的书写系统,任务与语言间的平行设计可实现对语言能力与问题解决能力的精细评估。评估结果显示,开源大模型在语法密集型任务和多文字系统语言上表现较差。我们进一步证明,MiLiC-Eval有助于推动低资源语言研究,提升对多样化书写系统处理及语言适应过程的理解。
原文摘要 · Abstract (English)
Large language models (LLMs) excel in high-resource languages but struggle with low-resource languages (LRLs), particularly those spoken by minority communities in China, such as Tibetan, Uyghur, Kazakh, and Mongolian. To systematically track the progress in these languages, we introduce MiLiC-Eval, a benchmark designed for minority languages in China, featuring 24K instances across 9 tasks. MiLiC-Eval focuses on underrepresented writing systems. Its parallelism between tasks and languages can provide a faithful and fine-grained assessment of linguistic and problem-solving skills. Our evaluation reveals that open-source LLMs perform poorly on syntax-intensive tasks and multi-script languages. We further demonstrate how MiLiC-Eval can help advance LRL research in handling diverse writing systems and understanding the process of language adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。