为智能代理设计可管控的自主机制,防止错误持续扩大。
Intelligence as Managed Autonomy: Failure, Escalation, and Governance for Agentic AI Systems
- 提出四层架构SMARt,通过状态切换管理认知漂移
- 理论证明系统在特定条件下能自动升级、限制输出、确保控制权回收
- 适用于医疗、机器人等高风险场景,支持安全扩展应用范围
随着自主智能体在机器人和人机环境中规模扩大,幻觉与持续但无依据的行为仍是未解难题。本文不将失败归因于模型或对齐缺陷,而是揭示无边界自主性的架构脆弱性——即假定智能体应无视不确定性持续运行。为此提出‘受控自主’理论:智能行为体现为检测知识漂移、暂停推理、尝试恢复,并在可靠性下降时主动放弃控制的能力。我们通过SMARt(自管理多层级自主推理,带受控/撤销转换)模型实现该理论,采用四层结构:稳定态、元认知态、辅助态与受控态。基于时序受保护的佩特里网建模,形式化证明系统具备可升降级、抑制无效输出、保障治理可达性的特性。进一步分析表明,在满足完备性与一致性前提下,引入领域特定触发集可系统性保障安全性。由于触发机制具有适应性,SMARt模型支持智能体操作范围的安全渐进扩展。结论指出:在自主生命周期中形式化失败管理,是实现可靠且受控人工智能的关键一步。
原文摘要 · Abstract (English)
As autonomous and agentic AI systems scale in robotic and human-machine environments, managing hallucination and persistent but unjustified action remains an open challenge. Rather than attributing these failures solely to model or alignment limitations, this paper explores the architectural vulnerability of unbounded autonomy - the presumption that an agent should continue operating regardless of rising uncertainty. It introduces a theory of managed autonomy that defines intelligent behavior through the formal capacity to detect epistemic drift, suspend reasoning, attempt recovery, and ultimately surrender control when reliability diminishes. We instantiate this theory via the SMARt (Self-Managing Multi-tier Autonomous Reasoning with Regulated/Revoked transitions) model, a four-layer framework featuring Stable, Meta-cognitive, Assisted, and Regulated states. By developing a timed, guarded Petri net formulation, we establish theoretically bounded properties for the system, demonstrating how architecture can formally mandate escalation, constrain invalid outputs, and ensure governance reachability under specified conditions. We further analyze how incorporating domain-specific trigger sets across varied operational settings (e.g., healthcare, robotics, etc.) can systematically preserve safety, assuming completeness and soundness criteria are met. Because these triggers are designed to be adaptive, the SMARt model accommodates the safe, controlled expansion of an agent's operational scope over time. We conclude that formalizing failure management within the autonomy lifecycle is a crucial step toward realizing reliable and governed artificial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。