用教育学标准评估大模型作文反馈质量,提升评分与修改效果
FeedEval: Pedagogically Aligned Evaluation of LLM-Generated Essay Feedback
- 设计三维度评估框架:具体性、有用性、正确性,专评大模型生成的作文反馈
- 在ASAP++数据集上与专家判断高度一致,筛选后反馈使评分模型性能更优
- 适合教育科技开发者和自动评分系统研究者使用,助力高质量反馈生成
超越数值评分,自动化作文评分研究日益关注生成具解释性和可操作性的高质量反馈。为降低专家标注成本,以往工作常依赖大模型生成的反馈来训练评估模型,但缺乏显式质量验证,导致噪声传播。为此,我们提出FeedEval,一个基于大模型的反馈评估框架,从教育学角度出发,从具体性、有用性和正确性三个维度评估大模型生成的作文反馈。FeedEval采用本研究构建的数据集训练的专用评估器,对多个反馈候选进行评估并筛选优质反馈用于下游任务。在ASAP++基准上的实验表明,FeedEval与人类专家判断高度一致;使用其过滤后的高质量反馈训练的评分模型表现更优。此外,利用小型大模型进行改写实验也显示,经FeedEval识别的高质量反馈能带来更有效的作文修订。代码与数据集已开源:https://github.com/BBeeChu/FeedEval.git。
原文摘要 · Abstract (English)
Going beyond the prediction of numerical scores, recent research in automated essay scoring has increasingly emphasized the generation of high-quality feedback that provides justification and actionable guidance. To mitigate the high cost of expert annotation, prior work has commonly relied on LLM-generated feedback to train essay assessment models. However, such feedback is often incorporated without explicit quality validation, resulting in the propagation of noise in downstream applications. To address this limitation, we propose FeedEval, an LLM-based framework for evaluating LLM-generated essay feedback along three pedagogically grounded dimensions: specificity, helpfulness, and validity. FeedEval employs dimension-specialized LLM evaluators trained on datasets curated in this study to assess multiple feedback candidates and select high-quality feedback for downstream use. Experiments on the ASAP++ benchmark show that FeedEval closely aligns with human expert judgments and that essay scoring models trained with FeedEval-filtered high-quality feedback achieve superior scoring performance. Furthermore, revision experiments using small LLMs show that the high-quality feedback identified by FeedEval leads to more effective essay revisions. We release our code and curated datasets at: https://github.com/BBeeChu/FeedEval.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。