提出动态安全模型与游戏化攻防平台,提升大模型防御实用性。
Gandalf the Red: Adaptive Security for LLMs
- 分离攻击者与正常用户,建模多步交互的动态安全机制
- 发现内置防御会降低可用性,即使不拦截请求
- 适合关注大模型安全与用户体验平衡的研究者
当前大语言模型应用中的提示攻击防御评估常忽略两个关键因素:攻击行为的动态性以及严格防御对合法用户的使用成本。我们提出D-SEC(动态安全效用威胁模型),明确区分攻击者与合法用户,建模多步交互,并将安全-效用关系表达为可优化形式。为进一步改进评估,我们设计Gandalf——一个众包、游戏化的红队平台,用于生成真实、自适应的攻击。利用Gandalf,我们收集并发布包含27.9万条提示攻击的数据集。结合良性用户数据,分析揭示了安全与效用之间的权衡关系,表明集成在LLM中的防御(如系统提示)即便不拦截请求,也会损害可用性。实验验证了限制应用领域、纵深防御和自适应防御是构建安全且可用的LLM应用的有效策略。
原文摘要 · Abstract (English)
Current evaluations of defenses against prompt attacks in large language model (LLM) applications often overlook two critical factors: the dynamic nature of adversarial behavior and the usability penalties imposed on legitimate users by restrictive defenses. We propose D-SEC (Dynamic Security Utility Threat Model), which explicitly separates attackers from legitimate users, models multi-step interactions, and expresses the security-utility in an optimizable form. We further address the shortcomings in existing evaluations by introducing Gandalf, a crowd-sourced, gamified red-teaming platform designed to generate realistic, adaptive attack. Using Gandalf, we collect and release a dataset of 279k prompt attacks. Complemented by benign user data, our analysis reveals the interplay between security and utility, showing that defenses integrated in the LLM (e.g., system prompts) can degrade usability even without blocking requests. We demonstrate that restricted application domains, defense-in-depth, and adaptive defenses are effective strategies for building secure and useful LLM applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。