arXiv:2510.22830cs.CLcs.LG2025-10被引 2

用生成式模型总结长作文,提升自动评分准确率。

Exploration of Summarization by Generative Language Models for Automated Scoring of Long Essays

  • 先用生成模型概括长作文,再用于评分
  • 评分一致性指标QWK从0.822升至0.8878
  • 适合需要处理长文本的自动评分场景

BERT及其变体广泛应用于自动评分,但其512词元的输入限制在长作文评分中表现不足。为此,本研究探索通过摘要与提示工程,利用生成式语言模型实现长作文的自动评分。实验结果显示,在Learning Agency Lab Automated Essay Scoring 2.0数据集上,评分一致性指标QWK从0.822提升至0.8878,显著提高了评分准确性。

原文摘要 · Abstract (English)

BERT and its variants are extensively explored for automated scoring. However, a limit of 512 tokens for these encoder-based models showed the deficiency in automated scoring of long essays. Thus, this research explores generative language models for automated scoring of long essays via summarization and prompting. The results revealed great improvement of scoring accuracy with QWK increased from 0.822 to 0.8878 for the Learning Agency Lab Automated Essay Scoring 2.0 dataset.

自动评分生成模型长文摘要

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。