分析890项研究,发现大模型评分能力受文本表述影响,且自回归架构表现显著落后。
Autoscoring Anticlimax: A Meta-analytic Understanding of AI's Short-answer Shortcomings and Wording Weaknesses
- 通过元分析建模,量化大模型在短答案评分中的表现差异
- 自回归模型平均比编码器模型差0.37,难度越低任务反而越难评分
- 揭示词汇表大小、分词方式对评分结果的敏感性及种族偏差风险
自动化短答案评分在大语言模型应用中仍落后于其他领域。本研究对890项系统性文献中的评分结果进行元分析,采用混合效应元回归模型,以二次加权卡帕系数(QWK)为传统效应量。结果显示,人类专家评分难度对大模型性能无显著统计影响;尤其值得注意的是,某些人类认为最简单的评分任务,对大模型而言反而是最难的。无论出于研究者实现问题还是源于自回归训练模式,解码器仅架构平均比编码器架构低0.37,差距显著。此外,我们还测量了分词器词汇量等技术因素的影响,发现其回报递减,可能因部分词元训练不足所致。研究呼吁系统设计应更前瞻地应对自回归模型的已知统计缺陷。最后,额外实验展示了表述和分词对评分结果的敏感性,并在高风险教育场景中揭示了模型存在的种族偏见。代码与数据已公开。
原文摘要 · Abstract (English)
Automated short-answer scoring lags other LLM applications. We meta-analyze 890 culminating results across a systematic review of LLM short-answer scoring studies, modeling the traditional effect size of Quadratic Weighted Kappa (QWK) with mixed effects metaregression. We quantitatively illustrate that that the level of difficulty for human experts to perform the task of scoring written work of children has no observed statistical effect on LLM performance. Particularly, we show that some scoring tasks measured as the easiest by human scorers were the hardest for LLMs. Whether by poor implementation by thoughtful researchers or patterns traceable to autoregressive training, on average decoder-only architectures underperform encoders by 0.37--a substantial difference in agreement with humans. Additionally, we measure the contributions of various aspects of LLM technology on successful scoring such as tokenizer vocabulary size, which exhibits diminishing returns--potentially due to undertrained tokens. Findings argue for systems design which better anticipates known statistical shortcomings of autoregressive models. Finally, we provide additional experiments to illustrate wording and tokenization sensitivity and bias elicitation in high-stakes education contexts, where LLMs demonstrate racial discrimination. Code and data for this study are available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。