arXiv:2603.17373cs.CL2026-03被引 5

评测AI导师在教学中的安全与有效性,发现越大模型越可能误导学生。

SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems

  • 构建涵盖11类风险的多学科评测基准SafeTutors
  • 77.8%对话轮次中出现教学失误,远超初始的17.7%
  • 提醒教育AI需兼顾教学逻辑与长期互动安全

大语言模型正快速应用于AI导师,但现有评估仅关注解题准确率和通用安全性,未能捕捉其在师生交互中是否同时具备教学有效性与安全性。我们认为,教学安全不同于传统LLM安全:主要风险并非生成有害内容,而是通过过早透露答案、强化错误认知、放弃渐进引导等方式悄然破坏学习过程。为系统研究此问题,我们提出SafeTutors基准,从数学、物理、化学三门学科出发,联合评估安全与教学效果。该基准基于学习科学文献构建了包含11类伤害维度和48个子风险的理论框架。实验发现所有模型均存在广泛危害,模型规模无法可靠缓解问题;多轮对话使教学失败率从17.7%飙升至77.8%。不同学科间危害表现差异显著,表明缓解措施需具备学科针对性;单轮“安全且有帮助”的结果可能掩盖长期互动中的系统性教学失效。

原文摘要 · Abstract (English)

Large language models are rapidly being deployed as AI tutors, yet current evaluation paradigms assess problem-solving accuracy and generic safety in isolation, failing to capture whether a model is simultaneously pedagogically effective and safe across student-tutor interaction. We argue that tutoring safety is fundamentally different from conventional LLM safety: the primary risk is not toxic content but the quiet erosion of learning through answer over-disclosure, misconception reinforcement, and the abdication of scaffolding. To systematically study this failure mode, we introduce SafeTutors, a benchmark that jointly evaluates safety and pedagogy across mathematics, physics, and chemistry. SafeTutors is organized around a theoretically grounded risk taxonomy comprising 11 harm dimensions and 48 sub-risks drawn from learning-science literature. We uncover that all models show broad harm; scale doesn't reliably help; and multi-turn dialogue worsens behavior, with pedagogical failures rising from 17.7% to 77.8%. Harms also vary by subject, so mitigations must be discipline-aware, and single-turn "safe/helpful" results can mask systematic tutor failure over extended interaction.

AI导师教学安全多轮对话教育AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。