arXiv:2510.09030cs.CL2025-10被引 4

让AI自己优化作文评分标准,提升与人工评分的一致性。

Automated Refinement of Essay Scoring Rubrics for Language Models via Reflect-and-Revise

  • AI通过反思自身评分理由和与人工评分的差异,迭代优化评分标准。
  • 在TOEFL11和ASAP数据集上,QWK最高提升0.47,优于人工制定的详细标准。
  • 即使初始标准简单,也能达到甚至超过人工制定标准的效果,适合教育评估场景。

大型语言模型(LLM)的表现高度依赖于输入提示。受提示优化领域的启发,本研究探索通过改进LLM用于自动作文评分(AES)的评分标准来提升性能。具体而言,我们的方法促使模型通过反思自身评分依据以及与人工评分在样本文本上的差异,实现评分标准的迭代优化。在TOEFL11和ASAP数据集上,使用GPT-4.1、Gemini-2.5-Pro和Qwen-3-Next-80B-A3B-Instruct进行实验,结果显示,二次加权卡帕系数(QWK)分别提升了最高0.19和0.47。值得注意的是,即使采用简单的初始评分标准,该方法仍能达到或优于使用详细人工制定的评分标准的效果。研究结果强调了在基于LLM的自动作文评分中,迭代式评分标准优化对提升与人工评分一致性的关键作用。

原文摘要 · Abstract (English)

The performance of Large Language Models (LLMs) is highly sensitive to the prompts they are given. Drawing inspiration from the field of prompt optimization, this study investigates the potential for enhancing Automated Essay Scoring (AES) by refining the scoring rubrics used by LLMs. Specifically, our approach prompts models to iteratively refine rubrics by reflecting on models' own scoring rationales and observed discrepancies with human scores on sample essays. Experiments on the TOEFL11 and ASAP datasets using GPT-4.1, Gemini-2.5-Pro, and Qwen-3-Next-80B-A3B-Instruct show Quadratic Weighted Kappa (QWK) improvements of up to 0.19 and 0.47, respectively. Notably, even with a simple initial rubric, our approach achieves comparable or better QWK than using detailed human-authored rubrics. Our findings highlight the importance of iterative rubric refinement in LLM-based AES to enhance alignment with human evaluations.

自动评分大模型提示优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。