arXiv:2605.26508q-fin.RMcs.AI2026-05被引 3

为自主AI设计可时间一致的反事实风险定价系统,让每步动作前就计算安全成本。

Foundations of a Time-Consistent Counterfactual Actuarial Runtime for Autonomous AI Agents

  • 以动作级保险为核心,用反事实风险代价替代事后责任赔偿。
  • 构建了路径分解归一化的边界势能机制,防止策略操纵。
  • 提供运行时预算保障,适合高风险自主系统设计者参考。

我们提出一种自主AI代理的基础运行时精算层,其中每个产生副作用的动作都需承担一个时间一致、基于反事实的風險費用,该费用以合同固定的安全部分为基准,并在显式承保边界内计算。该框架将单个动作的保险作为基本分析单元,用事前交易层取代事后年度责任覆盖。本文建立四项结构性结果:(i) 在选定安全默认映射和延续策略下,存在明确的反事实费用,但不唯一;(ii) 承保边界内具有无分裂性,使路径分解动作可归约为边界势能,其推论表明抗操控性与边界设计直接相关;(iii) 提出不可逆权限溢价,分为严格正的动作用层级成分,并给出集合级稳健资本增加的充要条件;(iv) 建立保守的运行时门控定理,将高概率费用包络转化为已执行动作的预算保证。本研究确立了后续多层扩展的基础:经验配套实例化为精算动作接口与权限边界实验;机制设计配套研究操作者激励与跨边界聚合;动态承保配套研究经验评级与审计回放校准。本文定义了原始合约、费用恒等式、边界内无套利结果及预算保证,构成后续各层依赖的数学基础。

原文摘要 · Abstract (English)

We propose a foundational runtime actuarial layer for autonomous AI agents in which every side-effect-bearing action carries a time-consistent, counterfactual risk toll computed against a contractually fixed safe default, inside an explicit underwriting boundary. The framework treats per-action insurance as the primary unit of analysis and replaces post-hoc annual liability cover with a pre-action transaction layer. The paper establishes four structural results: (i) a well-defined counterfactual toll under a chosen safe-default mapping and continuation policy, with explicit non-uniqueness; (ii) a no-splitting property within an underwriting boundary that telescopes path-decomposed actions into a boundary potential, with a corollary tying gaming-resistance to boundary design; (iii) an irreversible-authority premium, split into a strictly positive action-level component and an if-and-only-if characterisation of the set-level robust capital increase; and (iv) a conservative runtime gating theorem that translates high-probability toll envelopes into an executed-action budget guarantee. The result is the mathematical base layer for a broader program: an empirical companion instantiates the runtime through an Actuarial Action Interface and authority-frontier experiments; a mechanism-design companion studies strategic operator incentives and cross-boundary aggregation; and a dynamic-underwriting companion studies experience rating and audit-replay calibration. The present paper states the primitive contract, the toll identity, the within-boundary no-arbitrage result, and the budget guarantee on which those later layers depend.

AI安全反事实推理精算模型自主系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。