用自由能原理量化多智能体系统风险,实现透明安全治理
Free Energy Risk Metrics for Systemically Safe AI: Gatekeeping Multi-Agent Study
- 基于自由能原理构建可适应不同场景的累积风险暴露度量
- 低渗透率门控机制即可显著提升自动驾驶车队整体安全性
- 无需复杂世界模型或海量数据,适合实际系统安全治理
本文探讨自由能原理在智能体与多智能体系统风险度量中的基础作用。基于该原理提出一种灵活的累积风险暴露度量,区别于依赖海量数据或复杂世界模型的传统安全理论。本框架仅需利益相关方指定对系统结果的偏好,即可生成透明、可解释的风险治理与缓解规则。该框架自然融合世界模型与偏好模型的不确定性,支持认知与价值双重谦逊、简洁且面向未来的决策。我们在简化的自动驾驶环境中验证该方法:多辆自动驾驶汽车由门控代理实时评估其邻域集体安全风险,并在必要时干预各车策略。结果表明,即使门控代理渗透率较低,也能显著提升车队整体安全性,产生显著正外部性。
原文摘要 · Abstract (English)
We investigate the Free Energy Principle as a foundation for measuring risk in agentic and multi-agent systems. From these principles we introduce a Cumulative Risk Exposure metric that is flexible to differing contexts and needs. We contrast this to other popular theories for safe AI that hinge on massive amounts of data or describing arbitrarily complex world models. In our framework, stakeholders need only specify their preferences over system outcomes, providing straightforward and transparent decision rules for risk governance and mitigation. This framework naturally accounts for uncertainty in both world model and preference model, allowing for decision-making that is epistemically and axiologically humble, parsimonious, and future-proof. We demonstrate this novel approach in a simplified autonomous vehicle environment with multi-agent vehicles whose driving policies are mediated by gatekeepers that evaluate, in an online fashion, the risk to the collective safety in their neighborhood, and intervene through each vehicle's policy when appropriate. We show that the introduction of gatekeepers in an AV fleet, even at low penetration, can generate significant positive externalities in terms of increased system safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。