WrAFT用模块化设计实现论说文自动评分与精准反馈。
WrAFT: a Modularized Automated Writing Evaluation System for Argumentative Essays

- 分三模块:打分、表面反馈、深层反馈,提升可解释性。
- 评分准确率达QWK 0.84,RMSE仅0.44,接近人工水平。
- 适合语言教学、考试评估,支持免费在线使用。
本文提出WrAFT——一种模块化自动写作评估系统,可对论说文提供精准评分与有效反馈。系统将自动写作评估(AWE)任务拆分为评分、表层反馈和深层反馈三部分。研究中对比了LLaMA-3.3-70B-Instruct、GPT-4o与Claude 3.7等多种大语言模型,采用直接提示与监督微调两种方法。基于包含480篇TOEFL独立写作范文的自研数据集(附官方评分标准),基准测试显示,WrAFT在评分上达到顶尖性能:加权二次肯德尔系数(QWK)为0.84,均方根误差(RMSE)为0.44(评分范围0-5)。人工评估表明,系统生成的反馈认可度高:表层反馈通过率96.14%,深层宏观反馈93.03%,深层微观反馈94.69%。系统配备交互式界面,已公开免费使用。
原文摘要 · Abstract (English)
This study presents WrAFT, a Writing Assessment and Feedback Tool, that delivers both accurate and reliable scores and effective comprehensive feedback to argumentative essays. WrAFT adopts a modular design by dividing automated writing evaluation (AWE) tasks into scoring, surface-level feedback, and deep-level feedback. In building the system, various Large Language Models (LLMs) have been evaluated, including LLaMA-3.3-70B-Instruct, GPT-4o, and Claude 3.7, through both direct prompting and supervised fine-tuning approaches. A proprietary dataset of 480 TOEFL Independent Writing essays with official benchmark scores was utilized. Benchmark-based evaluation shows that WrAFT achieves state-of-the-art performance in scoring, with a quadratic weighted kappa (QWK) of 0.84 and a root mean square error (RMSE) of 0.44 against official scores on a scale of 0-5. Human evaluation of system-generated feedback also reveals high approval ratings: 96.14 percent for surface-level feedback, 93.03 percent for deep-level macro feedback, and 94.69 percent for deep-level micro feedback. An interactive user interface has been developed for the system and is publicly available and free to use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。