构建AI代理的外部基础设施,保障其交互安全与责任可追溯。
Infrastructure for AI Agents
- 设计外部协议系统来管理代理间的协作与责任归属
- 实现动作溯源、交互调控与有害行为检测三类核心功能
- 为未来智能代理生态提供可扩展的技术底座,适合政策与安全研究者
AI代理在开放环境中规划并执行交互任务,如OpenAI的Operator能使用浏览器完成商品比价与在线购买。当前研究多聚焦于直接修改代理行为以提升实用性与安全性,但此类方法难以应对异构代理间的复杂互动。因此,亟需外部协议与系统来规范其交互方式。例如,代理需更高效的通信协议以达成协作,且需能将行动归因于具体人类或其他法律实体,以建立信任并抑制滥用。为此,本文提出「代理基础设施」概念:即独立于代理之外、用于中介和影响其环境交互的技术系统与共享协议。正如互联网依赖HTTPS等协议,代理生态系统同样需要此类基础设施。我们识别出三大功能:1)对代理、用户或其他实体的动作、属性进行归因;2)调控代理间的交互行为;3)检测并纠正代理的有害行为。本文还列出了若干未完善的研究方向,涵盖应用场景分析、基础设施采纳、与现有互联网架构的关系、局限性及待解问题。推动代理基础设施的发展,有助于社会为更高级代理的广泛应用做好准备。
原文摘要 · Abstract (English)
AI agents plan and execute interactions in open-ended environments. For example, OpenAI's Operator can use a web browser to do product comparisons and buy online goods. Much research on making agents useful and safe focuses on directly modifying their behaviour, such as by training them to follow user instructions. Direct behavioural modifications are useful, but do not fully address how heterogeneous agents will interact with each other and other actors. Rather, we will need external protocols and systems to shape such interactions. For instance, agents will need more efficient protocols to communicate with each other and form agreements. Attributing an agent's actions to a particular human or other legal entity can help to establish trust, and also disincentivize misuse. Given this motivation, we propose the concept of \textbf{agent infrastructure}: technical systems and shared protocols external to agents that are designed to mediate and influence their interactions with and impacts on their environments. Just as the Internet relies on protocols like HTTPS, our work argues that agent infrastructure will be similarly indispensable to ecosystems of agents. We identify three functions for agent infrastructure: 1) attributing actions, properties, and other information to specific agents, their users, or other actors; 2) shaping agents' interactions; and 3) detecting and remedying harmful actions from agents. We provide an incomplete catalog of research directions for such functions. For each direction, we include analysis of use cases, infrastructure adoption, relationships to existing (internet) infrastructure, limitations, and open questions. Making progress on agent infrastructure can prepare society for the adoption of more advanced agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。