为AI代理事故建立分析框架,助力安全排查与预防。
Incident Analysis for AI Agents
- 从系统、上下文、认知三方面解析事故成因。
- 提出需保留活动日志、工具信息等关键数据。
- 适合安全研究者与部署方参考,提升风险应对能力。
随着AI代理的广泛应用,由其使用引发的事故(如提示注入导致数据泄露或未经授权购买)将日益增多。现有报告机制依赖公开数据,遗漏了代理思维链、浏览历史等敏感但关键的信息。为此,本文基于系统安全方法,提出三种事故成因:系统相关(如CBRN训练数据)、上下文相关(如提示注入)、认知相关(如误解用户请求)。并建议应保留活动日志、系统文档与访问权限记录、工具使用信息等,以辅助调查。同时提出报告内容规范与数据留存建议,帮助开发者和部署者在事故后提供有效支持。未来对代理事故的理解将成为管理风险的关键。
原文摘要 · Abstract (English)
As AI agents become more widely deployed, we are likely to see an increasing number of incidents: events involving AI agent use that directly or indirectly cause harm. For example, agents could be prompt-injected to exfiltrate private information or make unauthorized purchases. Structured information about such incidents (e.g., user prompts) can help us understand their causes and prevent future occurrences. However, existing incident reporting processes are not sufficient for understanding agent incidents. In particular, such processes are largely based on publicly available data, which excludes useful, but potentially sensitive, information such as an agent's chain of thought or browser history. To inform the development of new, emerging incident reporting processes, we propose an incident analysis framework for agents. Drawing on systems safety approaches, our framework proposes three types of factors that can cause incidents: system-related (e.g., CBRN training data), contextual (e.g., prompt injections), and cognitive (e.g., misunderstanding a user request). We also identify specific information that could help clarify which factors are relevant to a given incident: activity logs, system documentation and access, and information about the tools an agent uses. We provide recommendations for 1) what information incident reports should include and 2) what information developers and deployers should retain and make available to incident investigators upon request. As we transition to a world with more agents, understanding agent incidents will become increasingly crucial for managing risks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。