研究AI部署中控制权受限时的成本与策略,提出主权折扣概念。
Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the Model

- 定义'有限主权':对模型、数据、基础设施等层的部分控制权限。
- 实验显示完整日志提升故障诊断,预执行网关支持实时干预。
- 适合监管机构或企业用大模型却无模型所有权者参考。
AI控制研究关注如何在模型可能偏离目标的情况下安全部署,但多数协议假设部署者可控制模型及其运行环境。然而在受监管组织使用前沿模型的API或托管接口场景中,部署者通常仅能控制业务流程,无法获取模型权重、服务基础设施、内部日志或完整交互记录。本文提出‘有限主权’概念:在数据、模型、基础设施和交互层上拥有部分技术与合同访问权。研究表明,这些访问条件决定了实际可行的控制协议。论文构建了四层访问分类体系、协议-层级需求矩阵,并引入‘主权折扣成本’——即为弥补缺失访问而支付的合同、审计、架构调整、残余风险等代价。通过135万次合成案例仿真,在匿名化国家支付系统场景中验证结论。结果显示:完整日志提升诊断能力,预执行网关支持干预,日志与版本控制增强事后解释力,而范围限制虽提升安全性但降低实用性。因此,任何通用安全方案都应明确其访问假设。
原文摘要 · Abstract (English)
AI control research asks how to deploy models safely even when they may be misaligned, but many control protocols assume that the deployer can instrument the model and its surrounding pipeline. That assumption often fails for regulated organisations using frontier models through APIs or managed endpoints, where the deployer may control the business process but not the model weights, serving infrastructure, internal traces, update process, or full interaction logs. This paper introduces bounded sovereignty: partial technical and contractual access across the data, model, infrastructure, and interaction layers of the AI stack. It argues that these access conditions determine which control protocols can be executed in practice. The paper contributes a four-layer access typology, a protocol-by-layer requirements matrix, and the concept of sovereignty discount cost: the part of the control tax spent substituting for missing access through contracts, architecture, audit, vendor assurance, residual risk, or reduced system scope. It also reports a synthetic access-ablation experiment over 1.35 million synthetic case simulations and interprets the findings through an anonymised national-payments-infrastructure scenario. The experiment is not real-world payment-system evidence; it is a construct-validity exercise. The results show that complete logs improve diagnosis, a pre-execution gateway enables intervention, trace access and model-version control strengthen post-incident explanation, and scope restriction can improve safety while reducing usefulness. Control protocols proposed as general safety solutions should therefore state their access assumptions explicitly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。