提出可验证的智能体运行时架构,解决企业AI协作中的责任与能力分离问题。
A Contract-Centered Architecture for Scalable and Manageable Agentic Runtimes

- 用四个责任对象构建共享契约,明确能力、运行时、边界和数据的权责
- 提出可证伪的假设P1,量化能力与容量变化对系统响应的影响范围
- 设计随机交叉实验框架,支持四类结论判定,适合企业级AI治理团队使用
企业AI部署是跨业务单元、应用与AI团队、测试、平台工程、基础设施、安全、运维及数据治理的协调难题。现有用例基准仅能判断单个智能体是否完成任务,却无法明确能力、模型、运行机制、容量及企业数据的归属、变更、准入与证据应如何协同。本文提出四个共享组织契约:技能(可复用、版本化的能力与工作流资产)、支架(运行时编译器与监管者)、脚手架(执行/控制边界与非功能需求所有者),以及由首席信息官独立治理语义与遥测的栈外数据底座。运行时核心为A = <S, H, X>,数据底座位于栈外。核心贡献是一个有界、可证伪的假设P1:在指定运行区域,改变激活能力可保持容量-响应交互在预注册等效范围内;改变兼容的脚手架容量可保持能力语义不劣于原水平,且所需控制仍在声明的执行预算内。六个设计条件转化为可测量义务,其覆盖度、违规、不确定性、成本与排除项共同决定P1是否可判定。提出一种集群周期随机交叉实验(平衡顺序、重置/清洗、重复种子与故障场景、聚类感知不确定性),并设定四状态判决:支持、证伪、条件工程或不明确。本文贡献包括合同约束的运行时架构、源保持的数据底座及可证伪的测量协议。未报告已完成的实现、实验、数据集或实测结果。
原文摘要 · Abstract (English)
Enterprise AI deployment is a coordination problem across business units, application and AI teams, testing, platform engineering, infrastructure, security, operations, and data governance. Use-case benchmarks show whether one agent completes one task, but not how changing capabilities, models, runtime mechanisms, capacity, and enterprise data should be owned, changed, admitted, or evidenced together. We present four responsibility objects as shared organizational contracts: Skill (reusable, versioned capability and workflow asset), Harness (runtime compiler and governor), Scaffold (execution/control boundary and NFR owner), and a stack-external data substrate under independent CIO-governed semantics and telemetry. The runtime core is A = <S, H, X>, with the data substrate outside that stack. The central contribution is one bounded, falsifiable hypothesis, P1 (cost-aware capability-capacity separability): within a declared operating region, changing activated capability preserves the capacity-response interaction within a preregistered equivalence margin, while changing compatible Scaffold capacity preserves capability semantics up to a non-inferiority margin, and the required controls stay within a declared enforcement budget. Six design conditions become measured obligations whose coverage, violations, uncertainty, cost, and exclusions determine whether P1 is decidable. We propose a cluster-period randomized crossover experiment (balanced order, reset/washout, repeated seeds and failure regimes, cluster-aware uncertainty) with a four-state verdict: supported, falsified, conditional-engineering, or inconclusive. This paper contributes a contract-bounded runtime architecture, a source-preserving data substrate, and a falsifiable measurement protocol. It reports no completed implementation, experiment, dataset, or measured result.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。