首个连接写作与反馈的学术写作数据集,助力智能教学评估
Exposía: Teaching and Assessment of Academic Writing Skills for Research Project Proposals and Peer Feedback
- 构建了师生写作与互评数据集,覆盖撰写-反馈-修改全流程
- 发现不同大模型在提案评分与评语评分中表现各异,闭源模型更优
- 多维度联合评分提示策略效果最佳,适合课堂实际应用
我们提出 Exposía,首个公开的高等教育写作与反馈关联数据集,支持基于教育学原理的计算方法研究。该数据集来自计算机科学专业《科研入门》课程,包含学生研究提案、同伴与教师的评论与自由文本评语。数据完整反映学术写作的多阶段流程:起草、接收反馈、根据反馈修订。所有提案与评语均附有人工评估分数,采用我们设计的细粒度、教育学基础的评分框架。我们用 Exposía 基线测试当前最先进的大语言模型(LLMs)在两项任务上的表现:(1)自动评分提案,(2)自动评分学生评语。结果表明,两项任务需不同模型;闭源模型始终优于开源权重模型,推动开源模型性能提升研究。此外,将多个写作维度联合评分的提示策略效果最优,对课堂部署具有重要启示。
原文摘要 · Abstract (English)
We present Exposía, the first public dataset that connects writing and feedback in higher education, enabling research on educationally grounded computational approaches to teaching and evaluating academic writing. Exposía includes student research project proposals and peer and instructor feedback consisting of comments and free-text reviews. The dataset was collected in the "Introduction to Scientific Work" course of the Computer Science. Exposía reflects the multi-stage nature of the academic writing process that includes drafting, receiving feedback, and revising the writing based on the feedback received. Both the project proposals and peer feedback are accompanied by human assessment scores based on a fine-grained, pedagogically-grounded schema for writing and feedback assessment that we develop. We use Exposía to benchmark state-of-the-art large language models (LLMs) on two tasks: automated scoring of (1) the proposals and (2) the student reviews. We find that the two tasks benefit from different LLMs. Furthermore, closed-source models consistently outperform open-weight models, motivating further research on improving the performance of open-weight models preferred in classroom settings. Finally, we establish that a prompting strategy that scores multiple aspects of the writing together is the most effective, an important finding for classroom deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。