多语言教育场景下大模型表现受训练数据影响,低资源语言性能明显下降。
Multilingual Performance Biases of Large Language Models in Education
- 在8种语言上测试主流大模型的教育任务表现
- 低资源语言任务准确率显著低于英语,与训练数据量相关
- 建议使用前验证目标语言的适用性
大型语言模型(LLMs)正越来越多地应用于教育领域,但其应用仍以英语为主。本文评估了主流大模型在八种非英语语言(中文、印地语、阿拉伯语、德语、波斯语、泰卢固语、乌克兰语、捷克语)及英语中的四项教育任务表现:识别学生误解、提供针对性反馈、互动辅导和翻译评分。结果表明,模型性能与训练数据中该语言的占比呈正相关,低资源语言表现较差。尽管多数语言表现尚可,但相比英语普遍出现显著下降。因此,建议教育实践者在部署前先验证模型在目标语言上的适用性。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly being adopted in educational settings. These applications expand beyond English, though current LLMs remain primarily English-centric. In this work, we ascertain if their use in education settings in non-English languages is warranted. We evaluated the performance of popular LLMs on four educational tasks: identifying student misconceptions, providing targeted feedback, interactive tutoring, and grading translations in eight languages (Mandarin, Hindi, Arabic, German, Farsi, Telugu, Ukrainian, Czech) in addition to English. We find that the performance on these tasks somewhat corresponds to the amount of language represented in training data, with lower-resource languages having poorer task performance. Although the models perform reasonably well in most languages, the frequent performance drop from English is significant. Thus, we recommend that practitioners first verify that the LLM works well in the target language for their educational task before deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。