arXiv:2504.05716cs.LGcs.CY2025-04KDD被引 8

用大模型自动评学生反思,单智能体+少样本效果最好。

Single-Agent vs. Multi-Agent LLM Strategies for Automated Student Reflection Assessment

  • 用单智能体+少样本提示,将反思转为评分
  • 与人工评分匹配度最高,能更好识别学业风险生
  • 适合教育科技、智能评测方向研究者参考

我们探索使用大语言模型(LLMs)对开放式学生反思进行自动化评估,并预测学业表现。传统评估方法耗时且难以在教育场景中扩展。本研究采用两种评估策略(单智能体与多智能体)和两种提示技术(零样本与少样本),基于377名学生在三个学术学期中的5,278份反思数据进行实验。结果表明,单智能体结合少样本策略与人工评分的匹配率最高;同时,基于LLM评估的反思分数在识别学业风险学生和预测成绩任务中均优于基线模型。这些发现表明,LLMs可有效实现反思评估自动化,减轻教师负担,并及时为需要帮助的学生提供支持。本工作强调了将先进生成式AI技术融入教育实践以提升学生参与度和学业成效的潜力。

原文摘要 · Abstract (English)

We explore the use of Large Language Models (LLMs) for automated assessment of open-text student reflections and prediction of academic performance. Traditional methods for evaluating reflections are time-consuming and may not scale effectively in educational settings. In this work, we employ LLMs to transform student reflections into quantitative scores using two assessment strategies (single-agent and multi-agent) and two prompting techniques (zero-shot and few-shot). Our experiments, conducted on a dataset of 5,278 reflections from 377 students over three academic terms, demonstrate that the single-agent with few-shot strategy achieves the highest match rate with human evaluations. Furthermore, models utilizing LLM-assessed reflection scores outperform baselines in both at-risk student identification and grade prediction tasks. These findings suggest that LLMs can effectively automate reflection assessment, reduce educators' workload, and enable timely support for students who may need additional assistance. Our work emphasizes the potential of integrating advanced generative AI technologies into educational practices to enhance student engagement and academic success.

教育AI大模型评估反思分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。