36个大模型在提示注入攻击下超一半失败,参数规模和架构影响安全风险。
Systematically Analyzing Prompt Injection Vulnerabilities in Diverse LLM Architectures
- 通过144次测试分析不同模型对提示注入的响应机制。
- 56%的测试成功触发恶意行为,参数量与架构显著影响脆弱性。
- 发现攻击技术间存在关联,适合安全研究者和部署方参考。
本研究系统分析了36个大型语言模型在各类提示注入攻击下的脆弱性,该技术通过精心设计的提示诱导模型产生恶意行为。在144次提示注入测试中,观察到模型参数量与脆弱性存在强相关性,统计分析(如逻辑回归、随机森林特征分析)表明参数规模和架构显著影响其易受攻击程度。结果显示56%的测试成功触发了提示注入,凸显各类参数规模模型普遍存在安全隐患;聚类分析识别出与特定模型配置相关的独特脆弱性模式。此外,研究还发现某些提示注入技术之间存在关联,暗示潜在的共性漏洞。这些发现强调了在关键基础设施和敏感行业中部署大模型时,亟需建立多层、稳健的防御体系。成功的提示注入攻击可能导致数据泄露、未授权访问或虚假信息传播。未来研究应探索多语言、多步骤防御及自适应缓解策略,以增强大模型在多样化真实环境中的安全性。
原文摘要 · Abstract (English)
This study systematically analyzes the vulnerability of 36 large language models (LLMs) to various prompt injection attacks, a technique that leverages carefully crafted prompts to elicit malicious LLM behavior. Across 144 prompt injection tests, we observed a strong correlation between model parameters and vulnerability, with statistical analyses, such as logistic regression and random forest feature analysis, indicating that parameter size and architecture significantly influence susceptibility. Results revealed that 56 percent of tests led to successful prompt injections, emphasizing widespread vulnerability across various parameter sizes, with clustering analysis identifying distinct vulnerability profiles associated with specific model configurations. Additionally, our analysis uncovered correlations between certain prompt injection techniques, suggesting potential overlaps in vulnerabilities. These findings underscore the urgent need for robust, multi-layered defenses in LLMs deployed across critical infrastructure and sensitive industries. Successful prompt injection attacks could result in severe consequences, including data breaches, unauthorized access, or misinformation. Future research should explore multilingual and multi-step defenses alongside adaptive mitigation strategies to strengthen LLM security in diverse, real-world environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。