arXiv:2605.25632cs.AIcs.LG2026-05被引 4

为自主智能体的每项操作建立实时风险控制框架,防止意外损失。

Insuring Every Action: An Authority Frontier Framework for Runtime Actuarial Control of Autonomous AI Agents

论文配图:Insuring Every Action: An Authority Frontier Framework for Runtime Actuarial Control of Autonomous AI Agents
图 1 · 摘自论文原文
  • 通过固定默认值和资本预算,对每个动作进行实时定价与审批。
  • 在四个场景中验证,低预算下可完全避免损失,资本需求相差22倍。
  • 适合关注智能体安全、风险控制的研究者与开发者使用。

自主智能体频繁执行带有副作用的操作:数据库修改、退款、支付、外部承诺等。本文提出动作精算接口(AAI),一种确定性运行时合约,基于时间一致的风险映射,将每个动作与合同约定的安全默认值进行定价,并依据边界储备资本预算决定是否执行。进一步构建了权威前沿(Authority Frontier),用于衡量不同储备资本水平下释放的自主权限。该框架具备:(i) 具有费用上限的能力令牌的报价-绑定-提交协议;(ii) 将异构工具调用映射为可比权限单位的通用七类动作分类体系;(iii) 回放确定性与路径依赖的储备覆盖,采用alpha支出机制;(iv) 通过全储备需求C_full和Capital@k实现跨领域归一化。在四个代理环境中实例化AAI(数据库修改、客户服务退款、tau-bench零售与航空工具使用轨迹),并在一个实际Postgres面板中,三个托管于Azure的模型通过同一合约提出动作。前沿显示各领域普遍存在低储备拒绝、中度释放模式,仅在预算达到全储备需求时饱和;所需储备资本跨度达22倍(Capital@50从289到6457)。框架不强求领域统一形态,而是揭示各领域的精算几何特征。在真实面板中,合同在低预算下阻止了所有模型的实现实质损失,但面对拒绝时的承保持续性存在模型差异:模型身份成为精算承保变量。贡献在于提供了一个可基准化评估自主代理副作用实时精算控制的框架。

原文摘要 · Abstract (English)

Autonomous AI agents increasingly issue side-effect-bearing actions: database mutations, refunds, payments, external commitments. We propose the Actuarial Action Interface (AAI), a deterministic runtime contract that prices each such action against a contractually fixed safe default under a time-consistent risk mapping, and gates execution against a per-boundary reserve capital budget. We then develop the Authority Frontier, an evaluation primitive measuring how much autonomous authority the runtime releases at each level of reserve capital. The framework provides (i) a deterministic quote-bind-commit protocol with toll-bounded capability tokens; (ii) a universal seven-class action taxonomy mapping heterogeneous tool calls to comparable authority units; (iii) replay determinism and pathwise reserve coverage under alpha-spending; (iv) cross-domain normalization via full reserve demand C_full and capital metrics Capital@k. We instantiate AAI across four agentic environments (database mutation, customer-service refund, and the public tau-bench retail and airline tool-use traces) and report a live Postgres panel in which three Azure-hosted models propose actions through the same contract. The frontier exhibits a common low-reserve refusal and intermediate-release pattern across domains, with saturation only where the budget grid reaches full reserve demand; required reserve capital varies by 22x (Capital@50 from 289 to 6457). The framework does not force domains into the same shape; it surfaces each domain's actuarial geometry. In the live panel the contract prevents realized loss across all three models at low budget while differing in underwriting persistence under denial: model identity is an actuarial underwriting variable. The contribution is a benchmark-ready evaluation framework for runtime actuarial control of autonomous-agent side effects.

智能体安全风险控制精算框架自动化决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。