用大模型自动评估学生批判性思维的细分能力
Toward LLM-Supported Automated Assessment of Critical Thinking Subskills
- 基于真实作文任务设计评分标准,分解批判性思维为可测子技能
- 微调Llama 3.1 8B在可区分性强的子技能上表现最佳
- 适合教育评估研究者与智能辅导系统开发者参考
随着人工智能生成内容、虚假信息和算法说服的泛滥,批判性思维——即评估证据、识别不可靠主张并独立判断的能力——正成为关键的人类素养。通过及时评估与反馈发展批判性思维至关重要,但教育数据挖掘领域尚未广泛开展对批判性思维的定义、测量与支持研究。本文探索了测量批判性思维“子技能”的可行性。我们以学生撰写论证性文章这一真实任务为基础,依据既有的能力发展框架制定编码评分标准,并完成人工标注语料库。随后,评估三种自动化评分方法:零样本提示、少样本提示与监督微调,分别在三个大语言模型(GPT-5、Llama 3.1 8B、ModernBERT)上实现。结果表明,微调Llama 3.1 8B表现最优,尤其在具备明显可区分性、各水平标签分布均衡的子技能上表现突出;而在需要捕捉细微差异或标签不平衡的子技能上效果较差。本研究为在真实教育场景中实现批判性思维技能的规模化评估迈出初步一步。未来研究应继续结合自动化评估与人工验证,更准确地检测与衡量动态的高阶思维能力。
原文摘要 · Abstract (English)
As the world becomes increasingly saturated with AI-generated content, disinformation, and algorithmic persuasion, critical thinking - the capacity to evaluate evidence, detect unreliable claims, and exercise independent judgment - is becoming a defining human skill. Developing critical thinking skills through timely assessment and feedback is crucial; however, there has not been extensive work in educational data mining on defining, measuring, and supporting critical thinking. In this paper, we investigate the feasibility of measuring "subskills" that underlie critical thinking. We ground our work in an authentic task where students operationalize critical thinking by writing argumentative essays. We developed a coding rubric based on an established skills progression and completed human coding for a corpus of student essays. We then evaluated three distinct approaches to automated scoring: zero-shot prompting, few-shot prompting, and supervised fine-tuning, implemented across three large language models (GPT-5, Llama 3.1 8B, and ModernBERT). Fine-tuning Llama 3.1 8B achieved the best results and demonstrated particular strength on subskills with highly separable proficiency levels with balanced labels across levels, while lower performance was observed for subskills that required detection of subtle distinctions between proficiency levels or imbalanced labels. Our exploratory work represents an initial step toward scalable assessment of critical thinking skills across authentic educational contexts. Future research should continue to combine automated critical thinking assessment with human validation to more accurately detect and measure dynamic, higher-order thinking skills.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。