arXiv:2605.30151cs.AI2026-05

AI数学任务评估能力随版本更新波动,但少样本提示可显著提升性能。

Temporal Stability and Few-Shot Prompting in Math Task Assessment

  • 用少量示例任务引导模型,提升其对数学题认知需求的分类能力。
  • 新版模型表现不一:通用模型准确率稳定在58%,教育专用模型从75%降至50%。
  • 少样本提示让两者准确率均提升,通用模型达67%,专用模型恢复至75%。

随着AI工具日益融入教育场景,其长期稳定性与对提示工程的响应性成为关注焦点。本纵向研究考察了不同AI工具使用任务分析指南(TAG;Stein & Smith, 1998)对数学任务认知需求进行分类的能力,重点关注模型版本迭代与少样本提示(每类认知需求提供两个示例任务)的影响。测试对象为通用型模型Gemini与教育专用模型Coteach,二者在先前基准测试中表现优异。测试分三阶段进行:基线、版本更新后重测、引入少样本提示后重测。结果表明,仅更新模型版本效果参差:Gemini准确率保持58%,而Coteach从75%下降至50%。但采用少样本提示后,两模型性能均显著提升:Gemini达67%,Coteach恢复至75%。研究显示,提示工程的改进效果优于被动模型升级,且版本更新未必提升专业教育任务表现,对教育者选择和评估AI工具具有重要启示。

原文摘要 · Abstract (English)

As AI tools become increasingly integrated into educational contexts, questions arise about both their stability over time and their responsiveness to prompt engineering techniques. This longitudinal study focused on different AI tools' ability to use the Task Analysis Guide (TAG; Stein \& Smith, 1998) to classify the cognitive demand of mathematics tasks. In particular, it examined whether this classification ability changed with (1) model version updates over time and (2) few-shot prompting using exemplar tasks. We tested a general-purpose AI tool (Gemini) and an education-specific AI tool (Coteach). The specific tools were selected because of their relatively high performance on relevant published benchmarks and prior task-specific tests. Models were tested at baseline, retested with model version updates, and then tested again using few-shot prompting (two exemplar tasks for each cognitive demand category). Results revealed that newer model versions alone produced mixed effects: Gemini's accuracy remained stable at 58\%, while Coteach's accuracy decreased from 75\% to 50\%. However, few-shot prompting improved both models' performance: Gemini increased to 67\% and Coteach recovered to 75\% accuracy. These findings demonstrate that prompt engineering techniques can have larger and more reliable effects than passive model improvements, and that version updates may not always improve performance on specialized educational tasks. The study has important implications for how educators and researchers should approach AI tool selection, evaluation, and implementation in educational contexts.

AI评估少样本提示数学教育

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。