arXiv:2508.08629cs.CYcs.AI2025-08被引 5

梳理50种教育类大模型攻击,识别高危威胁并提出防护框架。

Securing Educational LLMs: A Generalised Taxonomy of Attacks on LLMs and DREAD Risk Assessment

  • 构建涵盖模型与基础设施的50种攻击分类体系
  • 识别出令牌窃取、对抗提示等4类关键高危攻击
  • 适合教育机构和开发者用于安全防护设计

由于效率高、生产力提升显著,教育领域正广泛采用大型语言模型(LLMs)。面向教师、学习者及机构的教育类大模型(eLLMs)正增强教学、学习与学术管理效能。然而其在教育环境中的部署带来严峻网络安全挑战。当前尚缺乏对LLMs攻击的全面图谱及其在教育场景的影响分析。本文提出一个涵盖50种攻击的通用分类体系,按攻击目标分为针对模型或其基础设施两类。基于DREAD风险评估框架,评估这些攻击在教育领域的严重性。结果表明,令牌走私、对抗性提示、直接注入及多步越狱是eLLMs的关键威胁。该分类体系、教育场景应用及风险评估,将帮助学术界与产业界构建更具韧性的防护方案,保障学习者与机构安全。

原文摘要 · Abstract (English)

Due to perceptions of efficiency and significant productivity gains, various organisations, including in education, are adopting Large Language Models (LLMs) into their workflows. Educator-facing, learner-facing, and institution-facing LLMs, collectively, Educational Large Language Models (eLLMs), complement and enhance the effectiveness of teaching, learning, and academic operations. However, their integration into an educational setting raises significant cybersecurity concerns. A comprehensive landscape of contemporary attacks on LLMs and their impact on the educational environment is missing. This study presents a generalised taxonomy of fifty attacks on LLMs, which are categorized as attacks targeting either models or their infrastructure. The severity of these attacks is evaluated in the educational sector using the DREAD risk assessment framework. Our risk assessment indicates that token smuggling, adversarial prompts, direct injection, and multi-step jailbreak are critical attacks on eLLMs. The proposed taxonomy, its application in the educational environment, and our risk assessment will help academic and industrial practitioners to build resilient solutions that protect learners and institutions.

大模型安全教育AI风险评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。