arXiv:2604.03976cs.AIcs.CE2026-04被引 6

为自主AI代理设计可量化的风险管理体系,保障用户在交易中的实际利益。

Quantifying Trust: Financial Risk Management for Trustworthy AI Agents

论文配图:Quantifying Trust: Financial Risk Management for Trustworthy AI Agents
图 1 · 摘自论文原文
  • 借鉴金融风控思路,构建端到端的代理风险标准(ARS)
  • 用户在任务失败时可获合同约定的赔偿,风险可控
  • 适合关注AI安全落地与责任归属的研究者与开发者

现有可信AI研究多聚焦模型内部属性,如偏见缓解、对抗鲁棒性与可解释性。随着AI系统演变为部署于开放环境的自主代理,并日益与支付或资产相连,信任的实质转向端到端结果:代理是否完成任务、遵循用户意图,以及避免造成实质性或心理伤害的故障。这类风险本质上是产品级的,无法仅靠技术防护消除,因代理行为具有固有随机性。为此,我们提出一种互补框架——基于风险管理的代理风险标准(Agentic Risk Standard, ARS)。ARS借鉴金融承保机制,将风险评估、承保与补偿整合进单一交易框架,实现对用户在代理交互中的保护。当发生执行失败、目标错位或意外结果时,用户可获得预先定义且合同可强制执行的赔偿。这使信任从对模型行为的隐含预期,转变为明确、可衡量、可执行的产品承诺。我们还通过仿真研究分析了将ARS应用于代理交易的社会效益。代码实现见 https://github.com/t54-labs/AgenticRiskStandard。

原文摘要 · Abstract (English)

Prior work on trustworthy AI emphasizes model-internal properties such as bias mitigation, adversarial robustness, and interpretability. As AI systems evolve into autonomous agents deployed in open environments and increasingly connected to payments or assets, the operational meaning of trust shifts to end-to-end outcomes: whether an agent completes tasks, follows user intent, and avoids failures that cause material or psychological harm. These risks are fundamentally product-level and cannot be eliminated by technical safeguards alone because agent behavior is inherently stochastic. To address this gap between model-level reliability and user-facing assurance, we propose a complementary framework based on risk management. Drawing inspiration from financial underwriting, we introduce the \textbf{Agentic Risk Standard (ARS)}, a payment settlement standard for AI-mediated transactions. ARS integrates risk assessment, underwriting, and compensation into a single transaction framework that protects users when interacting with agents. Under ARS, users receive predefined and contractually enforceable compensation in cases of execution failure, misalignment, or unintended outcomes. This shifts trust from an implicit expectation about model behavior to an explicit, measurable, and enforceable product guarantee. We also present a simulation study analyzing the social benefits of applying ARS to agentic transactions. ARS's implementation can be found at https://github.com/t54-labs/AgenticRiskStandard.

AI风险可信代理金融风控责任机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。