arXiv:2602.00979cs.CRcs.AI2026-02

提出细粒度攻击框架,暴露大模型评分系统漏洞

GradingAttack: Exposing Security Vulnerabilities in LLM Based Educational Grading Agents

  • 设计词级与提示级攻击,隐蔽操控评分结果
  • 提示攻击成功率更高,词级攻击更难被发现
  • 揭示现有评分系统缺乏防御能力,适合安全研究者

大型语言模型(LLMs)正被广泛应用于真实教育场景中的自动简答评分(ASAG),显著提升了评估效率与可扩展性。然而,当这些评分代理在实际环境中运行时,其易受对抗性攻击的特性引发了对其安全性和可信度的严重担忧。本文提出GradingAttack,一种细粒度的对抗攻击框架,系统评估基于LLM的教育评分代理的安全漏洞。具体而言,我们设计了词级和提示级攻击策略,在保持高隐蔽性的前提下操纵评分结果,暴露出当前代理部署中的根本缺陷。在多个数据集上的实验表明,两种攻击策略均能有效破坏评分代理,其中提示级攻击成功率更高,词级攻击则具备更强的隐蔽性。研究发现,当前基于LLM的教育代理普遍缺乏对抗攻击的防御能力,凸显了为关键教育应用开发安全可信系统的重要性和紧迫性。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed as educational agents for automatic short answer grading (ASAG) in real-world educational environments, significantly boosting assessment efficiency and scalability. However, when these grading agents operate ``in the wild'', their vulnerability to adversarial manipulation raises critical concerns about agent security and trustworthiness. In this paper, we introduce GradingAttack, a fine-grained adversarial attack framework that systematically evaluates the security vulnerabilities of LLM based educational grading agents. Specifically, we design token-level and prompt-level attack strategies that manipulate agent grading outcomes while maintaining high stealth, exposing fundamental weaknesses in current agent deployments. Experiments on multiple datasets demonstrate that both attack strategies effectively compromise grading agents, with prompt-level attacks achieving higher success rates and token-level attacks exhibiting superior stealth capability. Our findings reveal that current LLM based educational agents lack robust defenses against adversarial attacks, underscoring the urgent need for developing secure and trustworthy agent systems for critical educational applications.

大模型安全教育评分对抗攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。