首次系统评估主流大模型安全漏洞,提出有效防御框架。
Security Assessment and Mitigation Strategies for Large Language Models: A Comprehensive Defensive Framework
- 构建标准化评估框架,测试五类主流模型抗攻击能力。
- 漏洞率11.9%至29.8%,模型性能强弱与安全性无关。
- 开发高精度防御系统,检测准确率83%,误报仅5%。
大型语言模型正广泛应用于医疗、金融等关键基础设施,但其易受对抗性攻击的特性威胁系统完整性和用户安全。尽管部署日益广泛,却缺乏对主流LLM架构的全面安全评估,导致机构无法量化风险或为敏感应用选择更安全的模型。本研究填补该空白,建立标准化漏洞评估框架,并开发多层防御系统以应对识别出的威胁。我们系统评估了GPT-4、GPT-3.5 Turbo、Claude-3 Haiku、LLaMA-2-70B和Gemini-2.5-pro五类主流模型,针对10,000条涵盖六类攻击的对抗性提示进行测试。评估发现漏洞率在11.9%至29.8%之间,表明模型能力与安全性无相关性。为缓解风险,我们构建可投入生产的防御框架,在仅5%误报率下实现83%平均检测准确率。结果表明,系统化安全评估结合外部防御措施,是实现生产环境更安全大模型部署的有效路径。
原文摘要 · Abstract (English)
Large Language Models increasingly power critical infrastructure from healthcare to finance, yet their vulnerability to adversarial manipulation threatens system integrity and user safety. Despite growing deployment, no comprehensive comparative security assessment exists across major LLM architectures, leaving organizations unable to quantify risk or select appropriately secure LLMs for sensitive applications. This research addresses this gap by establishing a standardized vulnerability assessment framework and developing a multi-layered defensive system to protect against identified threats. We systematically evaluate five widely-deployed LLM families GPT-4, GPT-3.5 Turbo, Claude-3 Haiku, LLaMA-2-70B, and Gemini-2.5-pro against 10,000 adversarial prompts spanning six attack categories. Our assessment reveals critical security disparities, with vulnerability rates ranging from 11.9\% to 29.8\%, demonstrating that LLM capability does not correlate with security robustness. To mitigate these risks, we develop a production-ready defensive framework achieving 83\% average detection accuracy with only 5\% false positives. These results demonstrate that systematic security assessment combined with external defensive measures provides a viable path toward safer LLM deployment in production environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。