为持续演化的智能代理设定不可逾越的权限上限,确保授权安全可控。
Are You Still the Agent I Authorized? Earned Authority under a Fixed Ceiling for Evolving Agents
- 在授权时固定权限作用范围和最大影响上限,防止演化后权限失控。
- 证明了即使代理自我进化,其实际影响也不会超过用户设定的上限。
- 适用于长期运行、具备自主学习与任务调整能力的高风险智能代理系统。
长期运行的AI代理在部署后会持续演化:保留经验、获取新技能与工具、修改工作流程、分派任务并跨阶段迁移。这提升了适应性,却带来了独特的授权难题。具备工具调用能力的代理可能将模型错误或提示注入转化为实质性外部操作;当演化发生在有效授权期间,执行主体或上下文可能已不再符合用户最初评估的状态。演化可改变旧授权所能达成的效果范围及任务所需权限,权限可能上升、下降或变得不可比较。现有工具策略虽能限制行为,但无法判断授权是否应随变化延续。本文提出授权连续性问题:何时旧授权仍有效?活跃权限如何变化?何种边界不可突破?我们构建状态约束模型,在授权时刻固定一个转移容差区间与不可更改的影响上限。该区间决定授权是否在演化后存续;低于上限时,权限可自由收缩,仅在特定证据条件下才可扩展。我们区分请求效果与实际实现效果,并证明,在完全中介、合理效果抽象、弱化委派及监控完整性前提下,演化不会使受保护效果突破用户设定的上限。代理生成的证据可分配权限于上限之下,但无法提升上限。最后,我们将六类演化模式映射至其授权后果。
原文摘要 · Abstract (English)
Long-lived AI agents increasingly evolve after deployment by retaining experience, acquiring skills and tools, revising workflows, delegating work, and moving across task phases. This improves adaptation but creates a distinct authorization problem. Tool-enabled agents can turn model errors and prompt injections into consequential external actions; when evolution occurs under a live grant, the subject exercising that authority or the context in which it acts may no longer match what the user evaluated. Evolution can change both the effects reachable under an old grant and the authority required by the task, which may rise, fall, or become incomparable. Existing tool policies constrain actions but do not determine when a grant survives this change. We formulate authorization continuity: when does an existing grant remain valid, how may active authority change, and what boundary must never move? Our state-bound model fixes a transition envelope and an immutable effect ceiling at grant time. The envelope determines whether the grant survives a mutation; below the ceiling, authority may contract freely and expand only under specified evidence conditions. We distinguish requested from realized effects and prove that, under complete mediation, sound effect abstraction, attenuating delegation, and monitor integrity, mutation cannot amplify protected effects beyond the user-issued ceiling. Agent-produced evidence may allocate authority below the ceiling but cannot raise it. Finally, we map six mutation classes to their authorization consequences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。