简化评分标准也能保持作文自动评分准确率,更省计算资源。
Do We Need a Detailed Rubric for Automated Essay Scoring using Large Language Models?
- 用简化版评分标准替代详细标准,降低模型调用开销。
- 三种模型在简化标准下得分精度与详细标准相近,仅一模型表现下降。
- 适合关注效率的AI评分系统开发者参考,尤其多模型部署场景。
本研究探讨在基于大语言模型(LLM)的自动作文评分(AES)中,详细评分标准是否必要及其影响。尽管使用评分标准是标准做法,但创建详细标准需大量人力且增加令牌消耗。我们通过TOEFL11数据集,在四种LLM(Claude 3.5 Haiku、Gemini 1.5 Flash、GPT-4o-mini、Llama 3 70B Instruct)上比较了完整标准、简化标准和无标准三种条件下的评分准确性。结果显示,四款模型中有三款在简化标准下仍保持与详细标准相当的评分精度,同时显著减少令牌使用量;但其中一款模型(Gemini 1.5 Flash)在更详细标准下表现反而下降。结果表明,对多数基于LLM的AES应用而言,简化标准已足够,可在不牺牲精度的前提下提升效率。然而,不同模型性能差异明显,模型特定评估仍至关重要。
原文摘要 · Abstract (English)
This study investigates the necessity and impact of a detailed rubric in automated essay scoring (AES) using large language models (LLMs). While using rubrics are standard in LLM-based AES, creating detailed rubrics requires substantial ef-fort and increases token usage. We examined how different levels of rubric detail affect scoring accuracy across multiple LLMs using the TOEFL11 dataset. Our experiments compared three conditions: a full rubric, a simplified rubric, and no rubric, using four different LLMs (Claude 3.5 Haiku, Gemini 1.5 Flash, GPT-4o-mini, and Llama 3 70B Instruct). Results showed that three out of four models maintained similar scoring accuracy with the simplified rubric compared to the detailed one, while significantly reducing token usage. However, one model (Gemini 1.5 Flash) showed decreased performance with more detailed rubrics. The findings suggest that simplified rubrics may be sufficient for most LLM-based AES applications, offering a more efficient alternative without compromis-ing scoring accuracy. However, model-specific evaluation remains crucial as per-formance patterns vary across different LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。