arXiv:2604.22134cs.CL2026-04ACL被引 3

解决教育大模型被诱导直接给答案的问题,让辅导更安全且有效。

SHAPE: Unifying Safety, Helpfulness and Pedagogy for Educational LLMs

论文配图:SHAPE: Unifying Safety, Helpfulness and Pedagogy for Educational LLMs
图 1 · 摘自论文原文
  • 用知识掌握图建模教学逻辑,自动识别学习漏洞
  • 在9087组问答上测试,对抗攻击下安全提升显著
  • 适合教育AI研发者与评测人员参考

大型语言模型(LLMs)已广泛应用于教育场景。我们发现当前教育类LLM存在关键漏洞:学生可通过诱导性提问获取答案而非引导式指导,称为‘教学劫持’。为系统研究该问题,我们基于知识掌握图统一定义了安全、有用和教学性行为,并提出SHAPE基准,包含9,087个学生提问对,用于评估在对抗压力下的辅导表现。我们设计了一种图增强的辅导流程,通过查询推断前置概念,识别掌握缺口,并利用显式门控机制在指导与解题间动态切换。多模型实验表明,在两种教学劫持设置下,该方法显著提升安全性,同时在相同评估协议下保持接近天花板的帮助性。代码与数据已在https://github.com/MAPS-research/SHaPE公开。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have been widely explored in educational scenarios. We identify a critical vulnerability in current educational LLMs, pedagogical jailbreaks, where students use answer-inducing prompts to elicit solutions rather than scaffolded instructions. To enable systematic study, we unify and formalize safe, helpful, and pedagogical behaviors with a knowledge-mastery graph and introduce SHAPE, a benchmark of 9,087 student-question pairs for evaluating tutoring behavior under adversarial pressure. We propose a graph-augmented tutoring pipeline that infers prerequisite concepts from queries, identifies mastery gaps, and routes generation between instructing and problem-solving via explicit gating. Experiments across multiple LLMs show that our method yields significantly improved safety under two pedagogical jailbreak settings, while maintaining near-ceiling helpfulness under the same evaluation protocol. Our code and data are available at https://github.com/MAPS-research/SHaPE

教育AI安全评测知识图谱提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。