将语言学特征融入大模型,提升作文自动评分准确率。
Improve LLM-based Automatic Essay Scoring with Linguistic Features
- 用语言学特征增强大模型的作文评分能力
- 在同领域和跨领域提示下均优于基线模型
- 适合需要高精度评分的教育评估场景
自动作文评分(AES)可减轻教师批改负担。由于写作任务灵活多变,开发能应对多样题目生成的评分系统颇具挑战。现有方法分为两类:监督式特征方法与大语言模型(LLM)方法。前者性能较高但训练成本高;后者推理高效但准确率偏低。本文提出将语言学特征融入LLM-based评分框架,实验表明该混合方法在同域与跨域写作提示下均优于基线模型。
原文摘要 · Abstract (English)
Automatic Essay Scoring (AES) assigns scores to student essays, reducing the grading workload for instructors. Developing a scoring system capable of handling essays across diverse prompts is challenging due to the flexibility and diverse nature of the writing task. Existing methods typically fall into two categories: supervised feature-based approaches and large language model (LLM)-based methods. Supervised feature-based approaches often achieve higher performance but require resource-intensive training. In contrast, LLM-based methods are computationally efficient during inference but tend to suffer from lower performance. This paper combines these approaches by incorporating linguistic features into LLM-based scoring. Experimental results show that this hybrid method outperforms baseline models for both in-domain and out-of-domain writing prompts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。