arXiv:2604.12743cs.AI2026-04

测试AI能否提升数学题的认知难度,发现效果参差不齐。

Can AI Tools Transform Low-Demand Math Tasks? An Evaluation of Task Modification Capabilities

  • 用教师常用提示策略测试11款AI工具升级题目
  • 平均仅64%任务成功提升难度,最高88%最低33%
  • 专用工具优势不明显,且判断题难与改题难不相关

尽管已有研究探讨AI对数学任务质量的评估能力,但其提升现有低认知需求任务的能力仍不清楚。本研究检验了AI工具是否能有效升级低认知需求的数学任务。测试了11款工具,包括6款通用AI(如ChatGPT、Claude)和5款专为数学教师设计的工具(如Khanmigo、coteach.ai)。基于任务分析指南框架(Stein & Smith, 1998),对两类低需求任务进行提示修改,采用知识型教师可能采用的提示策略,而非优化寻找最优提示(即乐观典型结果)。结果显示,平均只有64%的任务被准确升级,各工具表现差异显著,从33%到88%不等。专用工具仅略优于通用工具。失败模式包括“未达标”(维持低认知需求)和“过度提升”(升至教师可能拒绝的过高目标类别)。有趣的是,任务分类准确率与任务升级成功率间存在微弱负相关(r = -.35),表明分类(评分)与生成式修改是两种不同能力。研究结果对理解AI在课程适配中的作用具有重要意义,并强调需发展专门方法支持教师修改教学材料。

原文摘要 · Abstract (English)

While recent research has explored AI tools' ability to classify the quality of mathematical tasks (arXiv:2603.03512), little is known about their capacity to increase the quality of existing tasks. This study investigated whether AI tools could successfully upgrade low-cognitive-demand mathematics tasks. Eleven tools were tested, including six broadly available, general-purpose AI tools (e.g., ChatGPT and Claude) and five tools specialized for mathematics teachers (e.g., Khanmigo, coteach.ai). Using the Task Analysis Guide framework (Stein & Smith, 1998), we prompted AI tools to modify two different types of low-demand mathematical tasks. The prompting strategy aimed to represent likely approaches taken by knowledgeable teachers, rather than extensive optimization to find a more effective prompt (i.e., an optimistic typical outcome). On average, AI tools were only moderately successful: tasks were accurately upgraded only 64% of the time, with different AI tool performance ranging from quite weak (33%) to broadly successful (88%). Specialized tools were only moderately more successful than general-purpose tools. Failure modes included both "undershooting" (maintaining low cognitive demand) and "overshooting" (elevating tasks to an overly ambitious target category that likely would be rejected by teachers). Interestingly, there was a small negative correlation (r = -.35) between whether a given AI tool was able to correctly classify the cognitive demand of tasks and whether the AI was able to upgrade tasks, showing that the ability to modify tasks (i.e., a generative task) represents a distinct capability from the ability to classify them (i.e., judgement using a rubric). These findings have important implications for understanding AI's potential role in curriculum adaptation and highlight the need for specialized approaches to support teachers in modifying instructional materials.

AI教育数学教学任务升级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。