arXiv:2505.14577cs.CL2025-05ACL被引 20

用大模型生成针对性评分特征,提升作文按特质打分的准确率。

TRATES: Trait-Specific Rubric-Assisted Cross-Prompt Essay Scoring

  • 基于评分标准生成特定维度特征,再融合通用与提示特征
  • 在主流数据集上所有特质评分均达新最好效果
  • 适合需要细粒度作文评估的研究者与教育应用

传统全自动作文评分研究历史悠久,但对按具体写作特质进行评估的关注不足。本文提出TRATES,一种基于评分标准、针对特定写作特质的跨提示作文评分框架。该框架利用大语言模型(LLM)根据特质评分标准生成特质相关特征(以评估问题形式表示),并据此评估作文。这些特质特征随后与通用写作质量特征及提示特定特征结合,训练一个简单的经典回归模型,用于预测未见提示下的作文特质得分。实验表明,TRATES在广泛使用的数据集上所有特质评分均达到新最佳性能,其中由LLM生成的特征贡献最为显著。

原文摘要 · Abstract (English)

Research on holistic Automated Essay Scoring (AES) is long-dated; yet, there is a notable lack of attention for assessing essays according to individual traits. In this work, we propose TRATES, a novel trait-specific and rubric-based cross-prompt AES framework that is generic yet specific to the underlying trait. The framework leverages a Large Language Model (LLM) that utilizes the trait grading rubrics to generate trait-specific features (represented by assessment questions), then assesses those features given an essay. The trait-specific features are eventually combined with generic writing-quality and prompt-specific features to train a simple classical regression model that predicts trait scores of essays from an unseen prompt. Experiments show that TRATES achieves a new state-of-the-art performance across all traits on a widely-used dataset, with the generated LLM-based features being the most significant.

作文评分大模型特质评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。