arXiv:2608.10153cs.AI2026-08

用四大学科理论构建企业智能体治理框架,解决多层失控难题

The CASE Framework: A Multi-Disciplinary Control Architecture for Governing Enterprise Agentic AI

  • 分层引入控制论、适应系统、人机协同与工程运维理论
  • 82%生产故障涉及多层耦合,35个公开部署均处最低成熟度
  • 提出非补偿性评估体系,适合监管合规与企业级智能体管理

企业部署自主智能体的速度远超治理能力,现有方法将单一学科(如面向确定性自动化的DevSecOps)强行应用于所有智能体层级。本文认为,智能体治理实为四个独立问题,各对应成熟学科:控制论用于个体智能体(意图为设定点,护栏为反馈,评估为观测);复杂自适应系统理论用于智能体群体(涌现使单体保证不可组合);监督控制论用于人机团队(所需多样性定律表明人类无辅助监督在结构上失败);工程运维扩展至集群(将错误预算延伸至决策质量,使自主性成为可控变量)。我们形式化各层,推导跨层耦合条件,揭示零接触部署悖论——某层卓越会挤压其他层。通过三组实证研究验证:82%已记录的生产故障为多层路径;22个生态工具中无一覆盖完整第二层(涌现);35个公开部署均位于最低成熟度等级。该现象称为‘涌现缺口’——在涌现层风险已暴露,但能力未配、实践缺失。提出五级成熟度模型,采用非补偿性瓶颈加权指数与评估工具,使CASE成为基于真实企业平台的科学治理模型。随着欧盟人工智能法案第14条要求有效人工监督,唯有满足必要多样性的架构才能实现真实而非形式化的监督。

原文摘要 · Abstract (English)

Enterprises are deploying autonomous AI agents faster than they can govern them, and prevailing approaches stretch a single discipline, typically DevSecOps built for deterministic automation, across every scale of agency. We argue that agentic AI governance is four problems, not one, each with a mature governing science. The CASE framework assigns Control theory to the individual agent (intent as setpoint, guardrails as feedback, evaluation as observation), complex Adaptive systems theory to agent collectives (where emergence makes single-agent assurance non-compositional), Supervisory cybernetics to human-agent teams (where the Law of Requisite Variety shows unaided human oversight fails structurally), and Engineering operations to fleets (extending error budgets to decision quality so autonomy becomes a controlled variable). We formalize each layer, derive cross-layer coupling conditions, including a zero-touch deployment paradox where excellence at one-layer strains the others, and trace twenty-plus enterprise controls to their classical constructs. Three empirical studies validate the thesis: 82 percent of documented production agent failures are multi-layer trajectories; none of 22 ecosystem tools offers full Layer 2 (emergence) coverage; and all 35 scored public deployments fall in the lowest maturity band. We name this mismatch, risk realized at the emergence layer against capability barely offered and practice absent, the Emergence Gap. A five-level maturity model with a non-compensatory bottleneck-weighted index and assessment instrument operationalizes CASE as a scientific rather than process maturity model, grounded in production enterprise agentic platforms. As EU AI Act Article 14 makes effective human oversight a legal requirement, only architectures satisfying requisite variety can make oversight real rather than ceremonial.

智能体治理多层架构企业AI合规监管

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。