用学习者自身写作特点评估LLM作文评分,更准发现弱项。
Towards Self-Referential Analytic Assessment: A Profile-Based Approach to L2 Writing Evaluation with LLMs

- 基于学习者个人写作档案,对比同一人各能力项表现
- LLM在识别弱项上优于单个真人评阅员
- 适合关注诊断性反馈而非排名的教育技术研究者
自动作文评分(AES)研究常依赖排名相关性指标验证分析性评分,但此类指标掩盖了写作能力内在维度间的关联及整体印象对分项评分的干扰。为此,本文提出一种自指式评估框架,聚焦于识别单个学习者的强弱项,而非跨学习者排名。我们在公开的ICNALE GRA数据集上开展实验,该数据集由最多80名训练有素的评阅员从整体和分析角度标注。通过双因子Rasch模型校准评阅者严苛度,获得十项分析维度与整体水平的公平平均分。在零样本条件下,比较人类评阅员与三种大语言模型(LLMs)的分析评分表现。结果表明,LLM在识别相对弱项(负面反馈)方面优于单个真人评阅员,而人类评阅员在识别相对强项(正面反馈)方面仍具优势。研究揭示了排名指标在分析性评估中的局限性,并证明了基于学习者档案的内生性方法在评估和部署LLM于AES中的价值。
原文摘要 · Abstract (English)
Automated essay scoring (AES) research often relies on rank-based correlation metrics to validate analytic assessment. However, such metrics obscure both intrinsic intercorrelations among analytic dimensions that arise from the structure of writing proficiency itself and halo effects, whereby holistic impressions bleed into fine-grained component scores. As a result, high correlations may mask a system's true diagnostic behaviour. In this study, we propose a novel self-referential assessment evaluation framework that focuses on identifying intra-learner strengths and weaknesses rather than assessing inter-learner rankings. We conduct experiments on the publicly available ICNALE GRA, a uniquely dense second-language writing dataset annotated holistically and analytically by up to 80 trained raters. To obtain reliable reference scores, we apply two-facet Rasch modelling to calibrate rater severity and derive fair average scores across ten analytic aspects and holistic proficiency. We compare the analytic scoring performance of human operational raters and three large language models (LLMs) in a zero-shot setting. Our results show that LLMs tend to outperform single human raters in identifying relative weaknesses (negative feedback) across several proficiency aspects, while human raters remain stronger at identifying relative strengths (positive feedback). Overall, our findings highlight the limitations of rank-based evaluation for analytic assessment and demonstrate the value of intra-learner, profile-based methods for assessing and deploying LLMs in AES.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。