arXiv:2507.19980cs.CL2025-07被引 4

用广义理论评估大模型作文评分可靠性,发现混合评分更优。

Exploring LLM Autoscoring Reliability in Large-Scale Writing Assessments Using Generalizability Theory

  • 采用广义理论分析人与大模型评分一致性
  • 大模型在故事叙述题中表现较稳定,整体可靠度接近人工
  • 人机混合评分显著提升可靠性,适合大规模考试

本研究基于广义理论,评估大语言模型(LLMs)在AP中文语言与文化考试写作题中的评分可靠性。选取两种题型:故事叙述和邮件回复,共14篇作文由两名受训人工评分员和七名AI评分员独立打分,每篇获得一个整体分和三个分项分(任务完成、表达、语言使用)。结果表明,尽管人工评分总体更可靠,但在特定条件下大模型表现合理一致,尤其在故事叙述任务中;结合人工与AI评分的复合评分方式显著提升了可靠性,支持混合评分模型在大规模写作评估中的应用价值。

原文摘要 · Abstract (English)

This study investigates the estimation of reliability for large language models (LLMs) in scoring writing tasks from the AP Chinese Language and Culture Exam. Using generalizability theory, the research evaluates and compares score consistency between human and AI raters across two types of AP Chinese free-response writing tasks: story narration and email response. These essays were independently scored by two trained human raters and seven AI raters. Each essay received four scores: one holistic score and three analytic scores corresponding to the domains of task completion, delivery, and language use. Results indicate that although human raters produced more reliable scores overall, LLMs demonstrated reasonable consistency under certain conditions, particularly for story narration tasks. Composite scoring that incorporates both human and AI raters improved reliability, which supports that hybrid scoring models may offer benefits for large-scale writing assessments.

大模型评分写作评估可靠性分析人机协作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。