为智能代理设计随情境动态生成的安全策略
Contextual Agent Security: A Policy for Every Purpose
- 根据任务上下文实时生成安全策略
- 支持人类可验证的即时决策
- 适用于多功能智能代理场景
判断一个行为的安全性需要了解其发生的上下文。对在多种情境中行动的人类而言,这显而易见:例如删除邮件是否合适,取决于邮件内容、目标(如清除敏感信息或清理垃圾)以及邮箱类型(工作或个人)。与人类不同,计算系统过去通常只在有限上下文中具备有限自主性,因此手动制定策略和用户确认(如手机应用权限或网络访问控制列表)虽不完美,但已足够限制有害操作。然而,随着通用代理(如自动化个人助理)的部署,我们需重新思考安全设计,以适应其复杂多变的上下文与能力。本文首次探索代理领域的上下文安全,提出上下文代理安全(Conseca)框架,实现即时生成、上下文相关且人类可验证的安全策略。
原文摘要 · Abstract (English)
Judging an action's safety requires knowledge of the context in which the action takes place. To human agents who act in various contexts, this may seem obvious: performing an action such as email deletion may or may not be appropriate depending on the email's content, the goal (e.g., to erase sensitive emails or to clean up trash), and the type of email address (e.g., work or personal). Unlike people, computational systems have often had only limited agency in limited contexts. Thus, manually crafted policies and user confirmation (e.g., smartphone app permissions or network access control lists), while imperfect, have sufficed to restrict harmful actions. However, with the upcoming deployment of generalist agents that support a multitude of tasks (e.g., an automated personal assistant), we argue that we must rethink security designs to adapt to the scale of contexts and capabilities of these systems. As a first step, this paper explores contextual security in the domain of agents and proposes contextual agent security (Conseca), a framework to generate just-in-time, contextual, and human-verifiable security policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。