arXiv:2508.14377cs.CLcs.AI2025-08

评测大模型对中文阅读难度的认知适配能力,发现其表现远低于预期。

ZPD-SCA: Unveiling the Blind Spots of LLMs in Assessing Students' Cognitive Abilities

  • 构建教师标注的中文阅读难度基准ZPD-SCA,覆盖不同学段。
  • 零样本下最强模型准确率仍低于随机猜测,上下文提示后提升近一倍。
  • 揭示大模型在教育认知判断中的系统性偏差,适合教育技术研究者参考。

大型语言模型(LLMs)在教育应用中展现出潜力,但其在评估阅读材料与学生认知发展水平匹配度方面的能力尚未充分探索。这一差距尤为关键,因为最近发展区(ZPD)理论强调学习资源需与学生认知能力(SCA)相匹配。尽管该匹配至关重要,目前仍缺乏针对中文教育背景下不同年龄组学生阅读理解难度的系统性研究。为此,我们提出ZPD-SCA,一个专为评估中文阅读理解难度阶段性的新基准,由60名特级教师标注,代表全国在岗教师前0.15%。实验结果表明,零样本场景下,Qwen-max和GLM的表现甚至低于随机猜测概率;而提供上下文示例后,部分模型准确率接近零样本基线的两倍。结果揭示了大模型在评估阅读难度方面的初步能力,也暴露其现有训练在教育适配判断上的局限。值得注意的是,即使最优模型仍存在系统性方向偏差,表明其难以精准匹配材料难度与学生认知水平。此外,不同文体间性能差异显著,凸显任务复杂性。我们期望ZPD-SCA能为评估与改进大模型在认知适配教育应用中的表现提供基础。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated potential in educational applications, yet their capacity to accurately assess the cognitive alignment of reading materials with students' developmental stages remains insufficiently explored. This gap is particularly critical given the foundational educational principle of the Zone of Proximal Development (ZPD), which emphasizes the need to match learning resources with Students' Cognitive Abilities (SCA). Despite the importance of this alignment, there is a notable absence of comprehensive studies investigating LLMs' ability to evaluate reading comprehension difficulty across different student age groups, especially in the context of Chinese language education. To fill this gap, we introduce ZPD-SCA, a novel benchmark specifically designed to assess stage-level Chinese reading comprehension difficulty. The benchmark is annotated by 60 Special Grade teachers, a group that represents the top 0.15% of all in-service teachers nationwide. Experimental results reveal that LLMs perform poorly in zero-shot learning scenarios, with Qwen-max and GLM even falling below the probability of random guessing. When provided with in-context examples, LLMs performance improves substantially, with some models achieving nearly double the accuracy of their zero-shot baselines. These results reveal that LLMs possess emerging abilities to assess reading difficulty, while also exposing limitations in their current training for educationally aligned judgment. Notably, even the best-performing models display systematic directional biases, suggesting difficulties in accurately aligning material difficulty with SCA. Furthermore, significant variations in model performance across different genres underscore the complexity of task. We envision that ZPD-SCA can provide a foundation for evaluating and improving LLMs in cognitively aligned educational applications.

大模型评估教育认知中文阅读

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。