arXiv:2506.04822cs.CL2025-06中稿 · AIED 2026被引 5

评估大模型在印尼真实教室手写作业中的自动评分与反馈效果

Evaluating Vision-Language and Large Language Models for Automated Student Assessment in Indonesian Classrooms

  • 用14000+份印尼四年级学生手写作业测试VLM和LLM
  • 手写识别差导致评分误差,但反馈仍具教学价值
  • 适合关注AI教育落地与跨语言评估的研究者

尽管视觉-语言模型(VLMs)和大型语言模型(LLMs)发展迅速,其在真实、资源匮乏课堂中实现人工智能驱动教育评估的有效性仍缺乏探索。本研究在覆盖数学与英语学科的14,000余份印尼四年级学生手写作业上评估了前沿VLM与LLM,这些作业符合当地国家课程标准。与以往基于整洁数字文本的研究不同,本数据集包含真实课堂中自然弯曲、风格多样的手写内容,带来真实的视觉与语言挑战。评估任务包括评分及依据评分量规生成个性化印尼语反馈。结果显示,VLM在手写识别方面表现不佳,引发后续LLM评分中的误差传播;然而,即使输入不完美,LLM生成的反馈仍具有教学实用性,揭示其在个性化与上下文相关性方面的局限。

原文摘要 · Abstract (English)

Despite rapid progress in vision-language and large language models (VLMs and LLMs), their effectiveness for AI-driven educational assessment in real-world, underrepresented classrooms remains largely unexplored. We evaluate state-of-the-art VLMs and LLMs on over 14K handwritten answers from grade-4 classrooms in Indonesia, covering Mathematics and English aligned with the local national curriculum. Unlike prior work on clean digital text, our dataset features naturally curly, diverse handwriting from real classrooms, posing realistic visual and linguistic challenges. Assessment tasks include grading and generating personalized Indonesian feedback guided by rubric-based evaluation. Results show that the VLM struggles with handwriting recognition, causing error propagation in LLM grading, yet LLM feedback remains pedagogically useful despite imperfect visual inputs, revealing limits in personalization and contextual relevance.

教育AI手写识别多模态评估低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。