arXiv:2507.14207cs.CRcs.AI2025-07

学生可伪造提示词绕过AI安全机制,研究提出检测工具

Mitigating Trojanized Prompt Chains in Educational LLM Use Cases: Experimental Findings and Detection Tool Design

  • 设计实验模拟学生提问,发现可诱导LLM生成不当内容
  • 在GPT-3.5和GPT-4中成功触发未授权输出,暴露安全漏洞
  • 开发原型工具TPG自动识别恶意提示链,适合教育AI开发者使用

大型语言模型(LLMs)在K–12教育中的应用既带来变革性机遇,也引发新风险。本研究探讨学生如何通过投毒式提示链诱导LLM生成不安全或非预期内容,从而绕过内置的内容安全防护机制。通过系统化实验,模拟真实教育场景中的多轮对话与教学查询,我们揭示了GPT-3.5与GPT-4的关键漏洞。本文详细呈现实验设计、关键发现,并开发原型检测工具TrojanPromptGuard(TPG),实现对投毒型教育提示的自动识别与缓解。研究成果旨在为人工智能安全研究者及教育技术从业者提供安全部署指导。

原文摘要 · Abstract (English)

The integration of Large Language Models (LLMs) in K--12 education offers both transformative opportunities and emerging risks. This study explores how students may Trojanize prompts to elicit unsafe or unintended outputs from LLMs, bypassing established content moderation systems with safety guardrils. Through a systematic experiment involving simulated K--12 queries and multi-turn dialogues, we expose key vulnerabilities in GPT-3.5 and GPT-4. This paper presents our experimental design, detailed findings, and a prototype tool, TrojanPromptGuard (TPG), to automatically detect and mitigate Trojanized educational prompts. These insights aim to inform both AI safety researchers and educational technologists on the safe deployment of LLMs for educators.

AI安全教育AI提示投毒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。