为雅思写作设计智能修订平台,自动评分并提供个性化反馈。
IELTS Writing Revision Platform with Automated Essay Scoring and Adaptive Feedback
- 用DistilBERT模型实现自动作文评分,支持自适应反馈。
- 平均误差0.66分,得分提升显著(均值+0.06分段,p=0.011)。
- 适合备考雅思者使用,尤其适合需针对性修改的考生。
本文提出并评估了一个针对雅思写作考试的智能修订平台。传统备考方式缺乏个性化反馈,且不贴合雅思评分标准。该平台采用直观用户界面、自动作文评分系统(AES)和基于雅思评分标准的精准反馈。系统将对话引导与写作界面分离,降低认知负荷,模拟真实考试环境。通过多轮设计研究(DBR),从规则基方法逐步优化至基于变压器的回归模型。早期阶段(2-3轮)显示规则方法存在中段压缩、准确率低及负R²问题。第4轮采用带回归头的DistilBERT模型,实现显著改进,平均绝对误差(MAE)为0.66,R²为正数。第5轮引入自适应反馈,结果显示得分提升显著(平均+0.060分段,p=0.011,Cohen's d=0.504),但效果因修改策略而异。研究建议自动化反馈应作为人工指导的补充,保守的表层修改比激进的结构调整更可靠。高分作文评估仍具挑战,未来需开展长期研究并邀请官方考官验证。
原文摘要 · Abstract (English)
This paper presents the design, development, and evaluation of a proposed revision platform assisting candidates for the International English Language Testing System (IELTS) writing exam. Traditional IELTS preparation methods lack personalised feedback, catered to the IELTS writing rubric. To address these shortcomings, the platform features an attractive user interface (UI), an Automated Essay Scoring system (AES), and targeted feedback tailored to candidates and the IELTS writing rubric. The platform architecture separates conversational guidance from a dedicated writing interface to reduce cognitive load and simulate exam conditions. Through iterative, Design-Based Research (DBR) cycles, the study progressed from rule-based to transformer-based with a regression head scoring, mounted with adaptive feedback. Early cycles (2-3) revealed fundamental limitations of rule-based approaches: mid-band compression, low accuracy, and negative $R^2$ values. DBR Cycle 4 implemented a DistilBERT transformer model with a regression head, yielding substantial improvements with MAE of 0.66 and positive $R^2$. This enabled Cycle 5's adaptive feedback implementation, which demonstrated statistically significant score improvements (mean +0.060 bands, p = 0.011, Cohen's d = 0.504), though effectiveness varied by revision strategy. Findings suggest automated feedback functions are most suited as a supplement to human instruction, with conservative surface-level corrections proving more reliable than aggressive structural interventions for IELTS preparation contexts. Challenges remain in assessing higher-band essays, and future work should incorporate longitudinal studies with real IELTS candidates and validation from official examiners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。