对比专家与大模型在议论文反馈中的差异,发现模型反馈更长但针对性弱。
FOXGLOVE: Understanding Goal-Oriented and Anchored Writing Feedback from Experts and LLMs on Argumentative Essays

- 构建专家与大模型统一标准的反馈数据集,涵盖69篇作文
- 模型反馈更复杂、提问少,且平均长度显著更长
- 尽管评分更高,但优势主要来自篇幅,非内容质量
尽管大语言模型(LLMs)越来越多地用于生成写作反馈,但缺乏对专家与模型在核心修订维度(目标导向性、锚定具体句子、优先级排序)上的系统性比较。本文提出FOXGLOVE数据集,包含696条由专业写作教师针对69篇高三年级议论文撰写的反馈,以及4个前沿大模型在相同协议下生成的1,644条反馈,共2,340条评论。对其中部分评论进行了专家质量评分。结果表明,教师与模型在反馈的目标分布和段落位置上相似,但在具体应反馈的句子上存在分歧;模型反馈更复杂、使用更少疑问句。此外,模型反馈在多数质量维度上得分更高,但这一优势主要源于评论更长。FOXGLOVE为系统比较人类与大模型反馈的异同提供了基础。
原文摘要 · Abstract (English)
While large language models (LLMs) are increasingly used to generate writing feedback, there remains no systematic comparison of LLM and expert feedback on the dimensions that writing research identifies as central to revision: goal-orientation, anchoring to specific sentences, and prioritization. We introduce FOXGLOVE, a dataset of 696 feedback comments written by trained writing instructors on 69 twelfth-grade argumentative essays, paired with 1,644 comments generated from four frontier LLMs under a shared protocol, totaling 2,340 comments. We provide expert quality ratings on a subset of both instructor and LLM comments. We find that instructors and LLMs distribute feedback similarly across goals and essay positions, yet instructors and models diverge on the specific sentences on which to provide feedback. Additionally, we find that models tend to write more complex feedback and use fewer questions than instructors. LLM feedback also receives higher ratings on most dimensions of quality, as rated by instructors, but much of this advantage appears to be attributable to lengthier comments. FOXGLOVE enables systematic comparison of where human and LLM feedback align, diverge, and differ.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。