为大模型设计可诊断认知错误的英语考试评测基准
From Test-taking to Cognitive Scaffolding: A Pedagogical Diagnostic Benchmark for LLMs on English Standardized Tests
- 将解题过程建模为认知路径,追踪思维轨迹
- 涵盖10576道题,29类任务,识别典型错误陷阱
- 适合教育AI研发者用于提升模型教学能力
随着大型语言模型(LLMs)在教育工具中的广泛应用,现有标准化测试评估多聚焦于二元准确率。然而,一个有效的AI导师应具备真实推理能力、清晰解题策略说明,并能诊断人类特定误解。为弥补这一差距,我们提出一种教学诊断框架,将英语标准化测试(EST)解题过程建模为认知框架中的遍历过程。基于此框架,我们构建了ESTBook,一个包含10,576道题目和29种任务类型的多模态基准,覆盖五大主流考试。与传统数据集不同,ESTBook通过形式化推理路径和干扰项合理化分析,深入刻画具体认知陷阱。大量实验证明,该诊断框架具有实用价值:识别认知路径有助于缩小性能差距,并通过引导式启发提升教学推理能力。
原文摘要 · Abstract (English)
As large language models (LLMs) are increasingly integrated into educational tools, current evaluations on standardized tests predominantly focus on binary outcome accuracy. Instead, an effective AI tutor must exhibit faithful reasoning, elucidate solution strategies, and diagnose specific human misconceptions. To bridge this gap, we introduce a pedagogical diagnostic framework that models English Standardized Test (EST) problem-solving as a traversal through a cognitive framework. Based on this framework, we present ESTBook, a multimodal benchmark encompassing 10,576 questions and 29 task types across five major exams. Unlike traditional datasets, ESTBook goes beyond data aggregation by enriching questions with formalized reasoning trajectories and distractor rationales that capture specific cognitive traps. Through extensive evaluations, we empirically demonstrate the practical utility of our diagnostic framework, showing that identifying cognitive trajectories facilitates the mitigation of performance gap and improves pedagogical reasoning through guided elicitation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。