对比大模型与人工评分理由,揭示打分差异原因
Comparison of Scoring Rationales Between Large Language Models and Human Raters
- 用GPT-4o、Gemini等大模型生成评分理由,对比人类评分逻辑
- 大模型评分准确率在测试中达到0.85以上(加权肯德尔系数)
- 适合关注AI评分可信度的研究者和教育评测人员
自动化评分的进步与机器学习和自然语言处理技术紧密相关。随着大语言模型(LLMs)的发展,使用ChatGPT、Gemini、Claude等生成式AI聊天机器人进行自动评分已受到关注。由于具备强推理能力,大模型可生成支持评分的理由。本研究通过大规模测试中的作文数据,基于二次加权肯德尔系数和归一化互信息评估GPT-4o、Gemini等大模型的评分准确性,并利用余弦相似度衡量评分理由之间的相似性。同时,基于理由嵌入向量的主成分分析揭示了理由的聚类模式。研究结果为理解大模型在自动评分中的‘思考’过程提供了洞见,有助于提升对人类评分与基于大模型的自动化评分背后推理机制的理解。
原文摘要 · Abstract (English)
Advances in automated scoring are closely aligned with advances in machine-learning and natural-language-processing techniques. With recent progress in large language models (LLMs), the use of ChatGPT, Gemini, Claude, and other generative-AI chatbots for automated scoring has been explored. Given their strong reasoning capabilities, LLMs can also produce rationales to support the scores they assign. Thus, evaluating the rationales provided by both human and LLM raters can help improve the understanding of the reasoning that each type of rater applies when assigning a score. This study investigates the rationales of human and LLM raters to identify potential causes of scoring inconsistency. Using essays from a large-scale test, the scoring accuracy of GPT-4o, Gemini, and other LLMs is examined based on quadratic weighted kappa and normalized mutual information. Cosine similarity is used to evaluate the similarity of the rationales provided. In addition, clustering patterns in rationales are explored using principal component analysis based on the embeddings of the rationales. The findings of this study provide insights into the accuracy and ``thinking'' of LLMs in automated scoring, helping to improve the understanding of the rationales behind both human scoring and LLM-based automated scoring.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。