提出可实时检测并缓解大模型三大安全威胁的统一框架
Unified Threat Detection and Mitigation Framework (UTDMF): Combating Prompt Injection, Deception, and Bias in Enterprise-Scale Transformers
- 基于自适应打补丁机制,实现多威胁联合检测
- 对提示注入攻击检测准确率达92%,欺骗输出减少65%
- 适合企业级AI系统部署,支持模型如Llama-3.1、GPT-4o
大型语言模型(LLMs)在企业系统中的快速应用暴露于提示注入攻击、策略性欺骗和偏见输出等漏洞,危及安全、信任与公平。在先前针对小型网络诱导欺骗率23.9%的对抗激活打补丁框架基础上,本文提出统一威胁检测与缓解框架(UTDMF),适用于企业级模型如Llama-3.1(405B)、GPT-4o和Claude-3.5。每模型开展700余次实验,结果表明:(1) 提示注入攻击(如越狱)检测准确率达92%;(2) 通过增强打补丁使欺骗输出减少65%;(3) 公平性指标(如人口统计偏差)提升78%。创新包括通用化打补丁算法用于多威胁检测、关于威胁交互的三项突破性假设(如企业工作流中的威胁链式传导),以及一个具备API接口的可部署工具包。
原文摘要 · Abstract (English)
The rapid adoption of large language models (LLMs) in enterprise systems exposes vulnerabilities to prompt injection attacks, strategic deception, and biased outputs, threatening security, trust, and fairness. Extending our adversarial activation patching framework (arXiv:2507.09406), which induced deception in toy networks at a 23.9% rate, we introduce the Unified Threat Detection and Mitigation Framework (UTDMF), a scalable, real-time pipeline for enterprise-grade models like Llama-3.1 (405B), GPT-4o, and Claude-3.5. Through 700+ experiments per model, UTDMF achieves: (1) 92% detection accuracy for prompt injection (e.g., jailbreaking); (2) 65% reduction in deceptive outputs via enhanced patching; and (3) 78% improvement in fairness metrics (e.g., demographic bias). Novel contributions include a generalized patching algorithm for multi-threat detection, three groundbreaking hypotheses on threat interactions (e.g., threat chaining in enterprise workflows), and a deployment-ready toolkit with APIs for enterprise integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。