用生成式摘要解决长作文评分中的信息丢失问题。
Cost-efficient generative AI summarization for scalable automated essay scoring in educational assessment

- 用GPT-5系列生成可控长度摘要,融合原文语言特征提升表征
- GPT-5 mini在评分一致性上表现最佳,QWK达0.812
- 适合关注可扩展写作评估与模型成本平衡的研究者
自动化作文评分(AES)能实现大规模评估和及时反馈,但受制于Transformer的输入长度限制,处理长作文时易造成信息丢失。本研究提出一种生成式AI辅助摘要框架,在保持评分可靠性的同时改进长作文表征。基于ASAP 2.0数据集,使用三种GPT-5变体(GPT-5、GPT-5 mini、GPT-5 nano)生成控制长度摘要,并作为下游AES模型的输入。为保留原始写作风格信号,将从完整作文中提取的手工语言特征与摘要表示融合,构建混合框架。评估涵盖评分性能、摘要质量与计算成本。评分可靠性以加权肯德尔协调系数(QWK)衡量,摘要质量通过词汇重叠、语义相似度、信息保留率与冗余度评估。结果表明,GPT-5 mini在人类评分一致性和总体表现上最优(QWK=0.812),GPT-5生成摘要质量最高;高分作文摘要质量下降,说明复杂文本更难压缩而不损失信息。研究揭示了模型容量、摘要保真度、成本效率与教育构念保留之间的权衡,为未来通用化研究提供了基线与消融实验参考。整体表明,生成式摘要在可扩展写作评估中具有潜力,但需谨慎验证信息保留与公平性。
原文摘要 · Abstract (English)
Automated essay scoring (AES) enables scalable assessment and timely feedback but remains challenged by transformer input-length limitations, which can cause information loss when processing long essays. This study proposes a generative AI-assisted summarization framework to improve long-form essay representation while maintaining scoring reliability. Using the ASAP 2.0 dataset, we generate controlled-length summaries with three GPT-5 variants (GPT-5, GPT-5 mini, and GPT-5 nano) and use them as inputs for downstream AES models. To preserve original writing signals, handcrafted linguistic features extracted from full essays are integrated with summary representations to form a hybrid framework. The approach is evaluated in terms of scoring performance, summarization quality, and computational cost. Scoring reliability is measured using quadratic weighted kappa (QWK), while summary quality is assessed through lexical overlap, semantic similarity, information retention, and redundancy metrics. Results show that GPT-5 mini achieves the highest agreement with human ratings, whereas GPT-5 produces the strongest summarization quality. Summary quality decreases for higher-scoring essays, indicating that more complex writing is more difficult to compress without information loss. These findings reveal trade-offs among model capacity, summary fidelity, cost efficiency, and preservation of educational constructs. This study provides an initial controlled evaluation of GPT-based summarization for AES and identifies important baselines and ablation studies required for future generalization. Overall, generative AI summarization offers a promising approach for scalable writing assessment while requiring careful validation of information preservation and fairness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。