构建首个高校作文反馈语料库,评估大模型生成反馈的准确性。
SEFORA: Student Essays with Feedback Corpus and LLM Feedback Evaluation Framework

- 构建师生互动式作文反馈数据集,含564篇多稿文本与8240条标注。
- 提出UniMatch评估框架,发现大模型生成反馈F1最高仅0.4。
- 揭示模型越生成越多反馈,越偏离教师实际关注重点。
有效的写作反馈是促进学生学习的关键因素,但规模化生成反馈成本高昂。大语言模型为扩展写作支持提供了自然路径,但存在两大障碍:缺乏公开数据集反映教师在真实课堂中的反馈方式;尚无可靠方法衡量生成反馈与教师实际反馈的一致性。本文提出SEFORA,一个公开语料库,包含564篇大学写作作业的多稿修订、8,240条教师段落级标注、任务提示、评分标准等信息,覆盖多种写作文体。同时提出UniMatch评估框架,用于开放式生成内容:将反馈切分为反馈单元,依据教师制定标准评分语义匹配度,并通过最优匹配计算可解释的精确率、召回率与F1。在74种不同配置下测试多个LLM,F1均未超过0.4,表明模型难以识别教师优先关注的反馈点,且生成量增加时性能持续下降。
原文摘要 · Abstract (English)
Effective writing feedback is among the strongest drivers of student learning, yet producing it at scale is labor-intensive. LLMs offer a natural path to scaling writing support, but two gaps stand in the way: few public corpora capture how instructors actually deliver feedback in real classrooms, and no reliable method measures whether generated feedback aligns with what an instructor would write. We address both. SEFORA is a public corpus pairing instructor inline feedback with assignment prompts, rubrics, scores, and multi-draft revisions across various college writing genres, comprising 564 drafts and 8,240 instructor annotations. UniMatch is a reference-based evaluation framework for open-ended generation: it segments feedback into feedback units, scores their semantic correspondence under instructor-derived criteria, and aligns them via optimal matching to yield interpretable precision, recall, and F1. Across 74 experimental configurations spanning multiple LLMs, no setting exceeds 0.4 F1. UniMatch reveals that models struggle to identify the feedback instructors would prioritize, and performance degrades as models generate more.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。