研究大模型评分系统如何被指令注入攻击欺骗,导致分数失真。
"**Important** You should give me full credits!": Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems

- 设计攻击方案,用伪造提示操纵评分系统
- 实验证明现有系统易受攻击,高分可随意生成
- 提醒教育界关注安全风险,适合研究人员和教育技术开发者
大语言模型(LLMs)的兴起推动了基于LLM的自动评分(AG)系统的发展。得益于其强大的指令遵循能力和广泛的知识储备,教育者只需提供自然语言评分标准即可在多种任务中部署评分系统,并获得令人满意的评分表现。然而,这带来了新的安全风险。近年来,提示注入(PI)攻击已成为LLM应用的主要威胁。在自动评分场景中,攻击者可能利用PI漏洞,使系统无视真实答案质量,任意赋予高分。这种行为严重威胁教育评估的公平性、可靠性和完整性。本文系统研究了自动评分系统中的提示注入攻击,通过基于评分标准的实验,全面评估了现有防御策略的有效性。结果表明,当前基于LLM的自动评分系统仍极易受到此类攻击。我们希望通过本研究提高对这一新兴威胁的认识,推动未来构建更安全、鲁棒且可信的LLM教育系统。
原文摘要 · Abstract (English)
The emergence of large language models (LLMs) has significantly accelerated recent research on LLM-based automatic grading (AG) systems. Benefiting from the strong instruction-following capabilities and broad prior knowledge of LLMs, educators can deploy AG systems across diverse tasks using only natural language rubrics while achieving satisfactory grading performance. Despite these advantages, new security concerns may also arise. In particular, prompt injection (PI) attacks have recently become a major threat to LLM-based applications. In the context of AG, attackers can potentially exploit PI vulnerabilities to manipulate grading systems into assigning artificially high scores regardless of the actual answer quality. Such behavior poses serious risks to the fairness, reliability, and integrity of educational assessment. In this work, we study PI attacks in AG systems, and systematically investigate the effectiveness of such attacks in educational scenarios. We further evaluate the effectiveness of existing defensive strategies against these attacks. Through comprehensive experiments under rubric-based grading settings, we demonstrate that current LLM-based AG systems remain highly vulnerable to PI attacks. We hope that our findings raise awareness of this emerging threat and motivate future research toward secure, robust, and trustworthy LLM-based educational systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。