统一生成作文评分与结构化反馈,提升评分一致性与可解释性。
A Unified Framework to Elicit Structured Feedback for Interpretable Multi-Trait Essay Scoring

- 先生成分层思维链反馈,再预测各维度和总分,实现评分与反馈统一
- 在中文数据集CFMS-34上,多维度评分准确率超基准模型12.3%
- 适合需要可解释评分的教育AI系统开发者使用
多特质自动作文评分需基于评分标准对相互关联的多个特质进行综合判断,而非孤立预测分数。现有增强反馈的方法常将反馈与评分分离或独立评估各特质,导致评分与反馈不一致且偏离评分标准。本文提出HiFTS,一个统一的自回归框架,在预测各特质分与总分前生成分层思维链(CoT)反馈。HiFTS从教师大模型中蒸馏出基于评分标准的分层思维链反馈,并训练学生模型联合生成反馈与分数。此外,采用分组相对策略优化(Group Relative Policy Optimization),结合综合奖励平衡分数一致性、校准性、反馈质量与结构有效性。推理时引入轻量级全局先验,减少长文本推理过程中的偏差。我们还构建了CFMS-34,一个包含951篇中文作文的多特质评分数据集,每篇均标注总分及34个基于评分标准的特质。在CFMS-34和ASAP++上的实验表明,HiFTS在保持强总分与特质分预测能力的同时,生成了连贯且符合评分标准的反馈。
原文摘要 · Abstract (English)
Multi-trait Automated Essay Scoring (AES) requires rubric-grounded reasoning across interdependent traits, rather than isolated score prediction. Existing feedback-enhanced methods often decouple feedback from scoring or assess traits independently, weakening score--feedback consistency and rubric alignment. We propose HiFTS, a unified autoregressive framework that generates hierarchical CoT feedback before predicting trait-level and holistic scores. HiFTS distills rubric-grounded hierarchical CoT feedback from a teacher LLM and trains student models to jointly generate feedback and scores. HiFTS further applies Group Relative Policy Optimization with a composite reward balancing score agreement, calibration, feedback quality, and structural validity. At inference, a lightweight global prior provides holistic guidance to reduce drift during long-form reasoning. We also introduce CFMS-34, a Chinese multi-trait AES dataset with 951 essays annotated with holistic scores and 34 rubric-based traits. Experiments on CFMS-34 and ASAP++ show that HiFTS achieves strong holistic and trait-level scoring while producing coherent, rubric-aligned feedback.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。