AI工具分类数学任务认知需求准确率仅63%,且易误判高低阶任务。
Baseline Performance of AI Tools in Classifying Cognitive Demand of Mathematical Tasks
- 用11个AI工具测试数学任务认知层级分类能力
- 平均准确率63%,最高未超83%
- 对极端任务类型(记忆/解题)易误判,偏好中等层级
教师面临时间压力,需在保持认知严谨性的同时适配个性化学习需求。本研究评估了11种AI工具(6种通用型:ChatGPT、Claude、DeepSeek、Gemini、Grok、Perplexity;5种教育专用型:Brisk、Coteach AI、Khanmigo、Magic School、School$.$AI)对数学任务认知需求的分类能力。采用基于研究的四层框架,评估其在无复杂提示下的表现。结果显示,平均准确率为63%,教育专用工具并未优于通用工具,且无一工具准确率超过83%。所有工具在处理认知需求极端任务(记忆与解题)时表现差,系统性倾向归类为中间层级(有/无联系的程序)。工具常给出看似合理但误导性的解释,尤其对新手教师具有说服力。错误分析显示,工具过度依赖表面文本特征,而非深层认知过程;其错误源于对多维度任务特征的错误推理,而非忽略关键维度。研究揭示了将AI融入教师备课流程的挑战,并强调需改进提示工程与教育专用工具开发。
原文摘要 · Abstract (English)
Teachers face increasing demands on their time, particularly in adapting mathematics curricula to meet individual student needs while maintaining cognitive rigor. This study evaluates whether AI tools can accurately classify the cognitive demand of mathematical tasks, which is important for creating or adapting tasks that support student learning. We tested eleven AI tools: six general-purpose (ChatGPT, Claude, DeepSeek, Gemini, Grok, Perplexity) and five education-specific (Brisk, Coteach AI, Khanmigo, Magic School, School$.$AI), on their ability to categorize mathematics tasks across four levels of cognitive demand using a research-based framework. The goal was to approximate the performance teachers will achieve with straightforward prompts. On average, AI tools accurately classified cognitive demand in only 63% of cases. Education-specific tools were not more accurate than general-purpose tools, and no tool exceeded 83% accuracy. All tools struggled with tasks at the extremes of cognitive demand (Memorization and Doing Mathematics), exhibiting a systematic bias toward middle-category levels (Procedures with/without Connections). The tools often gave plausible-sounding explanations likely to be persuasive to novice teachers. Error analysis of AI tools' misclassification of the broad level of cognitive demand (high vs. low) revealed that tools consistently overweighted surface textual features over underlying cognitive processes. Further, AI tools showed weaknesses in reasoning about factors that make tasks higher vs. lower cognitive demand. Errors stemmed not from ignoring relevant dimensions, but from incorrectly reasoning about multiple task aspects. These findings carry implications for AI integration into teacher planning workflows and highlight the need for improved prompt engineering and tool development for educational applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。