首个面向老挝语的多维度大模型评测基准,助力低资源语言公平评估。
LaoBench: A Large-Scale Multidimensional Lao Benchmark for Large Language Models
- 融合专家撰写与智能验证,构建高质量老挝语数据集
- 包含1.7万+样本,覆盖文化知识、教育和双语翻译三维度
- 支持黑盒评估,适合研究低资源语言模型的学者使用
大语言模型快速发展,但对低资源语言如老挝语的评估严重滞后。为此,我们推出首个大规模、高质量、多维度的老挝语评测基准LaoBench。该基准包含17,000+条专家精校样本,涵盖文化情境知识应用、中小学课程对齐教育任务以及老挝语、中文、英语间的双语翻译三个维度。数据集分为开源与预留两部分,预留部分可通过受控服务进行安全黑盒评估,提升评测公平性与数据安全性。LaoBench采用人机协作混合流程,结合专家撰写与智能代理验证,确保语言准确性、文化相关性和教育有效性。我们评估了多种前沿开源与闭源大模型,发现即使表现优异的多语言模型,在文化推理与翻译忠实度方面仍显著落后于人类专家。LaoBench旨在推动老挝语及其他东南亚低资源语言的模型研究,促进更包容的多语言评估体系。
原文摘要 · Abstract (English)
The rapid advancement of large language models (LLMs) has not been matched by their evaluation in low-resource languages, especially Southeast Asian languages like Lao. To fill this gap, we introduce \textbf{LaoBench}, the first large-scale, high-quality, and multidimensional benchmark for assessing LLM language understanding and reasoning in Lao. LaoBench contains \textbf{17,000+} expert-curated samples across three dimensions: culturally grounded knowledge application, curriculum-aligned K12 education, and bilingual translation among Lao, Chinese, and English. It includes open-source and held-out subsets, where the held-out portion enables secure black-box evaluation via a controlled service to improve fairness and data security. We construct LaoBench with a hybrid pipeline that combines expert authoring with agent-assisted verification, ensuring linguistic accuracy, cultural relevance, and educational validity. We evaluate diverse state-of-the-art open-source and closed-source LLMs, and find that even strong multilingual models lag behind human experts, particularly in culturally grounded reasoning and translation fidelity. We hope LaoBench will catalyze research on Lao and other underrepresented Southeast Asian languages for more inclusive multilingual evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。