arXiv:2602.14517cs.CLcs.LG2026-02中稿 · ITHET 2026

测试大模型在僧伽罗语和泰米尔语数学题上的表现,发现复杂题目效果明显下降。

Large Language Models for Math Education in Low-Resource Languages: A Study in Sinhala and Tamil

  • 用母语者独立编写三语数学题,避免翻译干扰
  • 基础算术跨语言表现好,复杂问题准确率显著降低
  • 提醒在非英语课堂使用AI前需做本地化评估

大型语言模型(LLMs)在数学推理上表现优异,正被用于教育辅导。但其在非英语语言,尤其是低资源语言中的可靠性仍不清楚。本文针对南亚学校广泛使用的僧伽罗语和泰米尔语开展研究,基于六类数学题型(从基础算术到复杂单位冲突与优化问题),评估四种主流大模型。为避免翻译误差混淆语言能力与翻译质量,构建了由母语者独立撰写的平行数据集,涵盖僧伽罗语、泰米尔语和英语版本。分析显示,基础算术推理在跨语言间表现稳定,而复杂推理任务在泰米尔语和僧伽罗语中出现显著退化。不同模型和题型的失败模式各异,表明英语中表现良好并不能保证多语言环境下可靠。研究结果对多语言课堂中部署AI工具具有直接指导意义,强调必须在非英语教育场景中进行语言特异性评估后再采用大模型作为数学辅导工具。

原文摘要 · Abstract (English)

Large language models (LLMs) have achieved strong results in mathematical reasoning, and are increasingly deployed as tutoring and learning support tools in educational settings. However, their reliability for students working in non-English languages, especially low-resource languages, remains poorly understood. We examine this gap by evaluating mathematical reasoning in Sinhala and Tamil -- two languages widely used in South Asian schools but underrepresented in artificial intelligence (AI) research. Using a taxonomy of six math problem types, from basic arithmetic to complex unit conflict and optimization problems, we evaluate four prominent large language models. To avoid translation artifacts that confound language ability with translation quality, we construct a parallel dataset in which each problem is independently authored in Sinhala and Tamil by native speakers, and in English by fluent speakers, all with strong mathematical backgrounds. Our analysis demonstrates that while basic arithmetic reasoning transfers robustly across languages, complex reasoning tasks show significant degradation in Tamil and Sinhala. The pattern of failures varies by model and problem type, suggesting that strong performance in English does not guarantee reliable performance across languages. These findings have direct implications for the deployment of AI tools in multilingual classrooms, and highlight the need for language-specific evaluation before adopting large language models as math tutoring aids in non-English educational contexts.

数学教育低资源语言大模型评估多语言AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。