为大模型对抗鲁棒性与合规性构建可验证的保障框架
Developing Assurance Cases for Adversarial Robustness and Regulatory Compliance in LLMs
- 分层部署架构+动态风险管理,防御越狱、随机化等攻击
- 结合欧盟人工智能法案,实现安全与合规双保障
- 适用于需高可靠性的医疗、金融等领域大模型应用
本文提出一种面向大语言模型(LLMs)对抗鲁棒性与监管合规性的保障案例构建方法。聚焦自然语言与代码语言任务,分析模型面临的漏洞,包括基于越狱、启发式及随机化的攻击。提出包含多阶段防护机制的分层框架,通过引入元层实现动态风险管理和推理,以应对模型漏洞的持续演化。通过两个典型保障案例,展示不同应用场景需采用定制化策略,确保AI系统在复杂威胁下仍具备稳健性与合规性。
原文摘要 · Abstract (English)
This paper presents an approach to developing assurance cases for adversarial robustness and regulatory compliance in large language models (LLMs). Focusing on both natural and code language tasks, we explore the vulnerabilities these models face, including adversarial attacks based on jailbreaking, heuristics, and randomization. We propose a layered framework incorporating guardrails at various stages of LLM deployment, aimed at mitigating these attacks and ensuring compliance with the EU AI Act. Our approach includes a meta-layer for dynamic risk management and reasoning, crucial for addressing the evolving nature of LLM vulnerabilities. We illustrate our method with two exemplary assurance cases, highlighting how different contexts demand tailored strategies to ensure robust and compliant AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。