AI辅助作文评分框架,人机协作提升大规模考试效率与质量
A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment

- 采用人机协同策略,对150-200字短文进行智能评分
- 模型评分与人工评分在多数维度达成中到高度一致
- 可识别需人工复核的关键案例,优化专家资源分配
将人工智能(尤其是大语言模型)融入教育评估,为提升评分效率和可扩展性提供了新机遇。本研究设计并验证了一种针对大规模国家级考试中书面作答的AI辅助评分框架。该方法聚焦约150-200字的短篇文本,采用人机协同策略,在保障评估质量的同时降低人工工作量。研究基于两次全国性考试的真实数据,每轮约5,000份学生作答。分析了模型评分与人工评分在多个评分维度上的一致性,以及决策流程对通过/不通过判定的影响。结果显示,模型与人工评价在多数维度上具中到高度一致性,支持该模式在该场景下的可行性。此外,所提出的修正工作流能有效识别出最需要人工干预的情况,实现专家精力的更高效配置。研究结论指出,唯有结合精心设计的人工监督,方可安全地将AI辅助评分集成至大规模评估体系。最后讨论了在国家评估系统中部署的实践意义,并提出未来研究方向,包括模型-人类对齐的长期监测及AI辅助评审流程可能引入的认知偏差分析。
原文摘要 · Abstract (English)
The integration of artificial intelligence (AI), particularly large language models (LLMs), into educational assessment has opened new opportunities to enhance the efficiency and scalability of grading processes. This study presents the design and validation of an AI-assisted scoring framework for written responses in a large-scale national assessment. The proposed approach focuses on short written texts of approximately 150-200 words and incorporates a human-in-the-loop strategy to preserve assessment quality while reducing manual workload. The study is grounded in a real operational context, using data from two recent editions of a nationwide test, each comprising approximately 5,000 student responses. We analyze the alignment between AI-generated scores and human raters across multiple rubric dimensions, as well as the impact of the proposed decision flow on pass/fail outcomes. Results show moderate to high agreement between the model and human evaluations in most dimensions, supporting the feasibility of AI assistance in this setting. Moreover, the proposed correction workflow identifies cases where human review is most valuable, enabling a more efficient allocation of expert effort. The findings suggest that AI-assisted scoring can be safely integrated into large-scale assessment processes only when combined with carefully designed human oversight. The paper concludes by discussing practical implications for deployment in national assessment systems and outlining future research directions, including longitudinal monitoring of model-human alignment and the analysis of potential cognitive bias introduced by AI-supported review workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。