arXiv:2508.05360cs.CYcs.AI2025-08被引 4

为教育类AI生成内容设计四重安全防护,保障5-16岁学生适用性。

Building Effective Safety Guardrails in AI Education Tools

  • 通过提示工程与课程对齐控制输出内容质量。
  • 引入独立异步内容审核机制,识别安全风险类别。
  • 坚持教师终审机制,确保教学内容安全可信。

生成式AI在教育领域快速发展,教师使用率上升,但其生成内容的安全性和适龄性引发关注。本文介绍英国政府首个公开可用的生成式AI工具——奥克国家学院的AI助教(Aila)所采取的安全防护措施。Aila旨在帮助教师规划符合国家课程、适合5-16岁学生的课程。为降低安全风险,实施四项关键防护:(1)提示工程确保输出符合教学法和课程要求;(2)输入威胁检测防范攻击;(3)独立异步内容审核代理(IACMA)评估输出是否符合预设安全类别;(4)采用人机协同模式,要求教师在课堂使用前审查内容。持续评估发现,需不断迭代优化防护策略,并推动跨领域协作,共享开源代码、数据集及经验。

原文摘要 · Abstract (English)

There has been rapid development in generative AI tools across the education sector, which in turn is leading to increased adoption by teachers. However, this raises concerns regarding the safety and age-appropriateness of the AI-generated content that is being created for use in classrooms. This paper explores Oak National Academy's approach to addressing these concerns within the development of the UK Government's first publicly available generative AI tool - our AI-powered lesson planning assistant (Aila). Aila is intended to support teachers planning national curriculum-aligned lessons that are appropriate for pupils aged 5-16 years. To mitigate safety risks associated with AI-generated content we have implemented four key safety guardrails - (1) prompt engineering to ensure AI outputs are generated within pedagogically sound and curriculum-aligned parameters, (2) input threat detection to mitigate attacks, (3) an Independent Asynchronous Content Moderation Agent (IACMA) to assess outputs against predefined safety categories, and (4) taking a human-in-the-loop approach, to encourage teachers to review generated content before it is used in the classroom. Through our on-going evaluation of these safety guardrails we have identified several challenges and opportunities to take into account when implementing and testing safety guardrails. This paper highlights ways to build more effective safety guardrails in generative AI education tools including the on-going iteration and refinement of guardrails, as well as enabling cross-sector collaboration through sharing both open-source code, datasets and learnings.

AI教育内容安全守护机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。