arXiv:2504.01253cs.CL2025-04被引 3

提升大模型评分准确率,解决部分正确答案误判问题。

Grade Guard: A Smart System for Short Answer Automated Grading

  • 通过调参与自反思机制增强大模型在短答评分中的任务专精度。
  • 引入不确定度评分,降低对模糊答案的误判率,最优模型降错率达23.64%。
  • 适合教育机构部署,尤其适用于需高精度、低人工干预的智能评分场景。

大型语言模型(LLM)在教育领域的应用推动了短答案自动评分(ASAG)的发展,显著提升了评分效率,缓解了人力短缺问题。然而,由于训练数据视角多样,现有方法在评估语义细微或部分正确的答案时仍易出错。为此,本文提出新框架 Grade Guard:1)通过均方根误差(RMSE)优化温度参数,提升模型任务专精度;2)引入不确定性评分(IS),使模型在输出分数的同时反映判断置信度;3)设计置信度感知损失(CAL),优化不确定性评分;4)引入基于优化后不确定度的自我反思机制,支持人工复核以减少错误评分。实验表明,Grade Guard 在 Upstage Solar Pro、Upstage Solar Mini、Gemini 1.5 Flash、GPT 4-o Mini 上分别比传统方法降低 19.16%、23.64%、4.00%、10.20% 的 RMSE。未来工作包括生成评分理由以提升可解释性,扩充带领域特性的标注数据集,并分析反馈以优化评分标准、减少偏见、实现个性化学习与多语言支持。

原文摘要 · Abstract (English)

The advent of large language models (LLMs) in the education sector has provided impetus to automate grading short answer questions. LLMs make evaluating short answers very efficient, thus addressing issues like staff shortage. However, in the task of Automated Short Answer Grading (ASAG), LLM responses are influenced by diverse perspectives in their training dataset, leading to inaccuracies in evaluating nuanced or partially correct answers. To address this challenge, we propose a novel framework, Grade Guard. 1. To enhance the task-based specialization of the LLMs, the temperature parameter has been fine-tuned using Root Mean Square Error (RMSE). 2. Unlike traditional approaches, LLMs in Grade Guard compute an Indecisiveness Score (IS) along with the grade to reflect uncertainty in predicted grades. 3. Introduced Confidence-Aware Loss (CAL) to generate an optimized Indecisiveness Score (IS). 4. To improve reliability, self-reflection based on the optimized IS has been introduced into the framework, enabling human re-evaluation to minimize incorrect grade assignments. Our experimentation shows that the best setting of Grade Guard outperforms traditional methods by 19.16% RMSE in Upstage Solar Pro, 23.64% RMSE in Upstage Solar Mini, 4.00% RMSE in Gemini 1.5 Flash, and 10.20% RMSE in GPT 4-o Mini. Future work includes improving interpretability by generating rationales for grades to enhance accuracy. Expanding benchmark datasets and annotating them with domain-specific nuances will enhance grading accuracy. Finally, analyzing feedback to enhance confidence in predicted grades, reduce biases, optimize grading criteria, and personalize learning while supporting multilingual grading systems will make the solution more accurate, adaptable, fair, and inclusive.

自动评分大模型教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。