研究AI系统临界点附近突变如何引发极端风险
Threshold Crossings as Tail Events for Catastrophic AI Risk
- 通过分析控制参数在临界点附近的随机波动
- 发现突发大规模跃迁概率与损伤尾部概率高度一致
- 为高风险AI系统的监控与控制提供新思路
我们分析了由分岔驱动的AI系统突变与潜在重尾结果分布之间的关联。通过研究控制参数在灾难性阈值附近的随机波动如何导致极端后果,我们揭示了在何种条件下,突然的大规模转变概率与最终损害分布的尾部概率高度吻合。该研究为监测、缓解和控制人工智能系统,在管理潜在灾难性人工智能风险方面提供了重要支持。
原文摘要 · Abstract (English)
We analyse circumstances in which bifurcation-driven jumps in AI systems are associated with emergent heavy-tailed outcome distributions. By analysing how a control parameter's random fluctuations near a catastrophic threshold generate extreme outcomes, we demonstrate in what circumstances the probability of a sudden, large-scale, transition aligns closely with the tail probability of the resulting damage distribution. Our results contribute to research in monitoring, mitigation and control of AI systems when seeking to manage potentially catastrophic AI risk.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。