arXiv:2604.09619cs.CYcs.AI2026-04被引 1

测试大模型在尼泊尔中小学教学中的适配性,发现其解释能力与文化契合度不足。

Assessing the Pedagogical Readiness of Large Language Models as AI Tutors in Low-Resource Contexts: A Case Study of Nepal's K-10 Curriculum

论文配图:Assessing the Pedagogical Readiness of Large Language Models as AI Tutors in Low-Resource Contexts: A Case Study of Nepal's K-10 Curriculum
图 1 · 摘自论文原文
  • 构建课程对齐评估框架,拆解教学有效性为七项二元指标
  • 顶尖模型整体可靠率达97%,但对低年级学生解释不清,简单题反而出错
  • 区域模型缺乏文化相关例证,建议人机协同部署并定制微调

将大语言模型(LLMs)融入教育生态有望实现个性化辅导的普惠化,但其在非西方、低资源环境下的适用性仍严重缺乏研究。本研究系统评估了四种前沿LLM——GPT-4o、Claude Sonnet 4、Qwen3-235B和Kimi K2——在尼泊尔小学五年级至十年级科学与数学教育中的教学潜力。我们提出一个新型课程对齐基准和细粒度评估框架,借鉴“自然语言单元测试”范式,将教学效能分解为七项二元指标:提示对齐、事实正确性、表达清晰度、上下文相关性、互动吸引力、有害内容规避及解题准确率。结果揭示显著的‘课程对齐差距’:尽管前沿模型(GPT-4o、Claude Sonnet 4)总体可靠性约达97%,但在教学清晰度与文化情境适配上存在明显缺陷。识别出两种普遍失败模式:‘专家诅咒’——模型能解难题却无法向初学者清晰讲解;‘基础谬误’——在更简单的低年级题目上表现反而下降,因无法适应青少年认知特点。此外,区域模型如Kimi K2在超过20%的交互中未能提供文化相关例证。研究指出,现成大模型尚不适合在尼泊尔课堂自主部署。我们提出‘人机协同’策略,并提供面向课程定制微调的方法论蓝图,以弥合全球AI能力与本地教育需求之间的鸿沟。

原文摘要 · Abstract (English)

The integration of Large Language Models (LLMs) into educational ecosystems promises to democratize access to personalized tutoring, yet the readiness of these systems for deployment in non-Western, low-resource contexts remains critically under-examined. This study presents a systematic evaluation of four state-of-the-art LLMs--GPT-4o, Claude Sonnet 4, Qwen3-235B, and Kimi K2--assessing their capacity to function as AI tutors within the specific curricular and cultural framework of Nepal's Grade 5-10 Science and Mathematics education. We introduce a novel, curriculum-aligned benchmark and a fine-grained evaluation framework inspired by the "natural language unit tests" paradigm, decomposing pedagogical efficacy into seven binary metrics: Prompt Alignment, Factual Correctness, Clarity, Contextual Relevance, Engagement, Harmful Content Avoidance, and Solution Accuracy. Our results reveal a stark "curriculum-alignment gap." While frontier models (GPT-4o, Claude Sonnet 4) achieve high aggregate reliability (approximately 97%), significant deficiencies persist in pedagogical clarity and cultural contextualization. We identify two pervasive failure modes: the "Expert's Curse," where models solve complex problems but fail to explain them clearly to novices, and the "Foundational Fallacy," where performance paradoxically degrades on simpler, lower-grade material due to an inability to adapt to younger learners' cognitive constraints. Furthermore, regional models like Kimi K2 exhibit a "Contextual Blindspot," failing to provide culturally relevant examples in over 20% of interactions. These findings suggest that off-the-shelf LLMs are not yet ready for autonomous deployment in Nepalese classrooms. We propose a "human-in-the-loop" deployment strategy and offer a methodological blueprint for curriculum-specific fine-tuning to align global AI capabilities with local educational needs.

AI助教教育AI低资源文化适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。