小模型物理推理常对但思路错,可能误导学生。
Dissecting Physics Reasoning in Small Language Models: A Multi-Dimensional Analysis from an Educational Perspective
- 用分阶段评估框架分析小模型解题思路
- 75%~98%正确答案含至少一处推理错误
- 适合教育AI安全评估与教学工具设计者
小型语言模型(SLMs)在教育场景中具有隐私和效率优势,但其可靠性依赖于多步推理能力。现有评测多关注最终答案正确率,忽视了‘答案对但过程错’的问题,可能强化学生误解。本文构建了包含3,162道高中及AP级物理题的Physbench数据集,题源为OpenStax,采用结构化参考解答并标注布鲁姆分类学层级,另含2,700组文化情境变体。通过P-REFS分阶段评估框架,对10个SLMs进行58,000次响应评估。结果表明:在最终答案正确的解法中,75%至98%存在至少一处推理错误;模型能力越强,失败模式从理解/建模转向执行阶段;高阶模型受情境变化影响小,中等模型性能显著下降。研究强调,教育AI安全需以推理保真度优先于答案正确性。
原文摘要 · Abstract (English)
Small Language Models (SLMs) offer privacy and efficiency for educational deployment, yet their utility depends on reliable multistep reasoning. Existing benchmarks often prioritize final answer accuracy, obscuring 'right answer, wrong procedure' failures that can reinforce student misconceptions. This work investigates SLM physics reasoning reliability, stage wise failure modes, and robustness under paired contextual variants. We introduce Physbench, comprising of 3,162 high school and AP level physics questions derived from OpenStax in a structured reference solution format with Bloom's Taxonomy annotations, plus 2,700 paired culturally contextualized variants. Using P-REFS, a stage wise evaluation rubric, we assess 10 SLMs across 58,000 responses. Results reveal substantial reliability gap: among final answer correct solutions, 75 to 98% contain at least one reasoning error. Failure modes shift with model capability; weaker models fail primarily at interpretation or modeling while stronger models often fail during execution. Paired contextual variations have minimal impact on top models but degrade the performance of mid-tier models. These findings demonstrate that safe educational AI requires evaluation paradigms that prioritize reasoning fidelity over final-answer correctness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。