对比机器学习与大模型在作文评分中的表现,发现各有优劣。
Operationalizing Automated Essay Scoring: A Human-Aware Approach
- 比较机器学习与大模型在评分中的表现差异
- 模型准确率高但解释性差,大模型解释更丰富
- 两者均存在偏见和对异常分数的鲁棒性不足
本文探讨了自动作文评分(AES)系统的人本化部署,关注准确性之外的关键维度。通过比较基于机器学习与大语言模型(LLMs)的方法,分析其在偏差、鲁棒性和可解释性方面的表现。研究发现,机器学习模型在评分准确率上优于大模型,但在可解释性方面较弱;而大模型虽提供更丰富的解释,却在偏差控制和对边缘分数的鲁棒性上表现不佳。本文揭示了不同方法间的权衡与挑战,为构建更可靠、可信的AES系统提供依据。
原文摘要 · Abstract (English)
This paper explores the human-centric operationalization of Automated Essay Scoring (AES) systems, addressing aspects beyond accuracy. We compare various machine learning-based approaches with Large Language Models (LLMs) approaches, identifying their strengths, similarities and differences. The study investigates key dimensions such as bias, robustness, and explainability, considered important for human-aware operationalization of AES systems. Our study shows that ML-based AES models outperform LLMs in accuracy but struggle with explainability, whereas LLMs provide richer explanations. We also found that both approaches struggle with bias and robustness to edge scores. By analyzing these dimensions, the paper aims to identify challenges and trade-offs between different methods, contributing to more reliable and trustworthy AES methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。