测试大模型判卷系统对提示注入攻击的脆弱性,发现小模型更易被攻破。
Adversarial Attacks on LLM-as-a-Judge Systems: Insights from Prompt Injections
- 区分内容作者攻击与系统提示攻击,系统化评估防御效果。
- 最高攻击成功率73.8%,小模型漏洞更明显,跨模型攻击成功率50.5%-62.6%。
- 适合关注大模型安全、评测系统鲁棒性的研究人员参考。
以大语言模型作为评判者来评估文本质量、代码正确性和论点强度的系统,容易受到提示注入攻击。我们提出一个框架,将内容作者攻击与系统提示攻击分离,并在四个任务上评估了五种模型(Gemma 3.27B、Gemma 3.4B、Llama 3.2 3B、GPT 4、Claude 3 Opus)在多种防御机制下的表现,每种条件下使用五十个提示进行测试。结果显示,攻击成功率最高达73.8%,较小模型更为脆弱,跨模型攻击的转移成功率在50.5%至62.6%之间。我们的结果与Universal Prompt Injection和AdvPrompter的研究存在差异。我们建议采用多模型委员会与对比评分机制,并已公开全部代码与数据集。
原文摘要 · Abstract (English)
LLM as judge systems used to assess text quality code correctness and argument strength are vulnerable to prompt injection attacks. We introduce a framework that separates content author attacks from system prompt attacks and evaluate five models Gemma 3.27B Gemma 3.4B Llama 3.2 3B GPT 4 and Claude 3 Opus on four tasks with various defenses using fifty prompts per condition. Attacks achieved up to seventy three point eight percent success smaller models proved more vulnerable and transferability ranged from fifty point five to sixty two point six percent. Our results contrast with Universal Prompt Injection and AdvPrompter We recommend multi model committees and comparative scoring and release all code and datasets
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。