arXiv:2504.05736cs.CLcs.AI2025-04被引 8

用排序引导打分,提升大模型中文作文评分能力

Rank-Then-Score: Enhancing Large Language Models for Automated Essay Scoring

  • 先用特征增强数据训练排序模型,再结合文章内容打分
  • 在HSK和ASAP数据集上,平均QWK均优于直接提示法
  • 特别适合中文作文自动评分,效果领先现有方法

近年来,大语言模型在多种任务中取得显著成果,但在自动作文评分(AES)领域潜力尚未充分挖掘。尤其与英文数据相比,中文作文评分方法发展滞后。本文提出一种基于大语言模型的细调框架——排序后打分(Rank-Then-Score, RTS)。具体而言,首先使用特征增强数据微调排序模型(Ranker),生成候选分数集;随后将该分数集与作文内容一并输入打分模型(Scorer),输出最终评分。在两个基准数据集HSK和ASAP上的实验表明,RTS在所有大模型和数据集上平均QWK均优于直接提示法(Vanilla),并在使用HSK数据集进行中文作文评分时达到最佳表现。

原文摘要 · Abstract (English)

In recent years, large language models (LLMs) achieve remarkable success across a variety of tasks. However, their potential in the domain of Automated Essay Scoring (AES) remains largely underexplored. Moreover, compared to English data, the methods for Chinese AES is not well developed. In this paper, we propose Rank-Then-Score (RTS), a fine-tuning framework based on large language models to enhance their essay scoring capabilities. Specifically, we fine-tune the ranking model (Ranker) with feature-enriched data, and then feed the output of the ranking model, in the form of a candidate score set, with the essay content into the scoring model (Scorer) to produce the final score. Experimental results on two benchmark datasets, HSK and ASAP, demonstrate that RTS consistently outperforms the direct prompting (Vanilla) method in terms of average QWK across all LLMs and datasets, and achieves the best performance on Chinese essay scoring using the HSK dataset.

作文评分大模型中文NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。