arXiv:2504.02323cs.CL2025-04被引 2

CoTAL让AI更懂跨学科作业评分,还能自动优化提示词

CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback

  • 结合证据中心设计与人机协作,自动构建评分系统
  • 相比基础模型,跨学科评分准确率最高提升38.9%
  • 适合教育AI研发者及需自动化评分的教师

大语言模型为教学辅助带来新可能。尽管已有研究探索教育场景中的提示工程,但其在科学、计算、工程等多领域的泛化能力仍不明确。本文提出链式思维提示+主动学习(CoTAL)方法,通过证据中心设计对齐评估与课程目标,结合人机协同提示工程实现响应评分自动化,并利用链式思维提示与师生反馈迭代优化题目、评分标准与提示词。实验表明,CoTAL使GPT-4在跨领域评分中表现显著提升,最高比无提示工程基线提高38.9%。教师与学生均认为该方法能有效评分并解释答案,其反馈进一步提升了评分准确性和解释质量。

原文摘要 · Abstract (English)

Large language models (LLMs) have created new opportunities to assist teachers and support student learning. While researchers have explored various prompt engineering approaches in educational contexts, the degree to which these approaches generalize across domains--such as science, computing, and engineering--remains underexplored. In this paper, we introduce Chain-of-Thought Prompting + Active Learning (CoTAL), an LLM-based approach to formative assessment scoring that (1) leverages Evidence-Centered Design (ECD) to align assessments and rubrics with curriculum goals, (2) applies human-in-the-loop prompt engineering to automate response scoring, and (3) incorporates chain-of-thought (CoT) prompting and teacher and student feedback to iteratively refine questions, rubrics, and LLM prompts. Our findings demonstrate that CoTAL improves GPT-4's scoring performance across domains, achieving gains of up to 38.9% over a non-prompt-engineered baseline (i.e., without labeled examples, chain-of-thought prompting, or iterative refinement). Teachers and students judge CoTAL to be effective at scoring and explaining responses, and their feedback produces valuable insights that enhance grading accuracy and explanation quality.

教育AI提示工程自动评分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。