用强化学习提升多维度作文评分准确率,让模型更懂人工评分标准。
Autoregressive Multi-trait Essay Scoring via Reinforcement Learning with Scoring-aware Multiple Rewards
- 设计基于QWK的多奖励机制,将评分标准融入训练过程
- 在多个数据集上显著提升低质量提示的评分表现
- 适合需要精准反馈的教育评估系统开发者
近年来,自动作文评分(AES)逐渐转向多维度评分以提供更丰富的反馈。与传统系统类似,多维度AES使用加权肯德尔协调系数(QWK)衡量与人工评分的一致性,贴合评分体系;然而其不可导特性阻碍了在神经网络中的直接应用。本文提出评分感知多奖励强化学习(SaMRL),通过设计基于QWK且包含均方误差惩罚的奖励函数,将真实评价方案融入训练过程。现有强化学习在AES中多限于分类模型,因需概率分布而性能下降;本文采用自回归得分生成框架,利用标记生成概率实现稳健的多维度评分预测。实证分析表明,SaMRL有效促进模型训练,显著提升以往表现较差提示的评分能力。
原文摘要 · Abstract (English)
Recent advances in automated essay scoring (AES) have shifted towards evaluating multiple traits to provide enriched feedback. Like typical AES systems, multi-trait AES employs the quadratic weighted kappa (QWK) to measure agreement with human raters, aligning closely with the rating schema; however, its non-differentiable nature prevents its direct use in neural network training. In this paper, we propose Scoring-aware Multi-reward Reinforcement Learning (SaMRL), which integrates actual evaluation schemes into the training process by designing QWK-based rewards with a mean-squared error penalty for multi-trait AES. Existing reinforcement learning (RL) applications in AES are limited to classification models despite associated performance degradation, as RL requires probability distributions; instead, we adopt an autoregressive score generation framework to leverage token generation probabilities for robust multi-trait score predictions. Empirical analyses demonstrate that SaMRL facilitates model training, notably enhancing scoring of previously inferior prompts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。