arXiv:2608.02024cs.AIcs.CY2026-08

构建教育场景下的LLM安全评估框架,发现动态对话中风险更高

EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers

论文配图:EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers
图 1 · 摘自论文原文
  • 结合教学场景与课程概念生成针对性对抗性交互
  • 动态多轮对话中模型漏洞更明显,安全防护不足
  • 适合教育AI开发者与政策制定者参考

大型语言模型(LLMs)在中小学教育中应用日益广泛,但现有安全评估很少关注模型与师生互动中出现的有害或不适当内容。为此,我们提出EduZone,一个面向多样化教育场景的LLM安全评估框架。该框架系统整合(1)学生与教师使用场景、(2)细粒度课程概念,以及(3)涵盖常规与教育特有风险的6类28个子类,生成情境化对抗性交互。交互在三种设置下构建:单轮请求、静态多轮对话和动态多轮对话。基于这些交互,我们对十款LLM在四个安全层级(拒绝、安全协助、带安全引导的风险协助、完全风险协助)下进行评估。结果表明,教育特定风险及动态多轮交互显著增加模型脆弱性,现有安全护栏未能有效应对。EduZone通过提供自动化、可扩展的评估框架,推动了中小学教育中LLM的安全发展。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used across diverse tasks in K-12 education, yet existing safety evaluations rarely examine how harmful or inappropriate content appears in interactions between LLMs and students or teachers. To address this, we present EduZone, an evaluation framework for LLM safety across diverse educational scenarios. Our framework systematically combines (1) student- and teacher-facing LLM usage contexts, (2) fine-grained curriculum concepts, and (3) 6 risk categories and 28 subcategories spanning both conventional and education-specific harms to generate contextually grounded adversarial interactions. We construct these interactions in three settings: single-turn requests, static multi-turn conversations, and dynamic multi-turn conversations. Using these interactions, we evaluate ten LLMs using four safety levels: refusal, safe assistance, risky assistance with safety guidance, and fully risky assistance. Our results reveal greater vulnerability to education-specific risks and dynamic multi-turn interactions, while existing safety guardrails fail to adequately address these risks. EduZone advances LLM safety in education by providing an automated, scalable evaluation framework that supports the development and deployment of safer LLMs in K-12 education.

LLM安全教育AI评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。