为大模型代理增强安全防护,提升企业级部署可靠性
Operationalizing CaMeL: Strengthening LLM Defenses for Enterprise Deployment
- 引入输入筛查与输出审计,防范提示注入和指令泄露
- 设计分层风险访问机制,兼顾安全性与使用体验
- 采用可验证中间语言,实现形式化安全保障
CaMeL(机器学习能力)提出基于能力的沙箱以缓解大型语言模型(LLM)代理中的提示注入攻击。尽管有效,但其假设用户输入可信、忽略侧信道威胁,并因双模型设计带来性能开销。本文识别上述问题,提出工程改进:(1) 初始输入的提示筛查,(2) 输出审计以检测指令泄露,(3) 分层风险访问模型以平衡可用性与控制,(4) 可验证的中间语言以提供形式化保证。这些升级使CaMeL更符合企业安全最佳实践,支持规模化部署。
原文摘要 · Abstract (English)
CaMeL (Capabilities for Machine Learning) introduces a capability-based sandbox to mitigate prompt injection attacks in large language model (LLM) agents. While effective, CaMeL assumes a trusted user prompt, omits side-channel concerns, and incurs performance tradeoffs due to its dual-LLM design. This response identifies these issues and proposes engineering improvements to expand CaMeL's threat coverage and operational usability. We introduce: (1) prompt screening for initial inputs, (2) output auditing to detect instruction leakage, (3) a tiered-risk access model to balance usability and control, and (4) a verified intermediate language for formal guarantees. Together, these upgrades align CaMeL with best practices in enterprise security and support scalable deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。