首个通用安全框架,让AI在保持高性能的同时大幅降低风险。
Domain-Agnostic Scalable AI Safety Ensuring Framework
- 用概率约束优化+安全分类模型,从机制上保障AI行为安全
- 强化学习中每千万动作仅3次碰撞,比基线减少99%以上
- 适合自动驾驶、医疗等高风险场景的AI部署,为安全提供理论支撑
随着AI系统在现实应用中日益普及,其安全性已成为关键挑战。本文提出首个领域无关的AI安全保障框架,兼具强安全保证与高性能,基于严格的理论基础。框架包含:(1)带概率约束的优化组件,(2)安全分类模型,(3)内部测试数据,(4)保守测试流程,(5)可量化的数据质量指标,(6)可计算梯度的连续近似损失函数。此外,首次在AI安全研究中建立量化规模定律,揭示数据量与安全-性能权衡的关系。在强化学习、自然语言生成和生产规划任务中的实验验证了框架有效性。特别地,在强化学习场景中,实现每1000万次动作仅3次碰撞,相较PPO-Lag基线(1000–3000次碰撞)在相当性能水平下显著提升安全表现,是以往方法无法达到的安全级别。本工作为安全关键领域的可信AI部署奠定了新基础。
原文摘要 · Abstract (English)
AI safety has emerged as a critical priority as these systems are increasingly deployed in real-world applications. We propose the first domain-agnostic AI safety ensuring framework that achieves strong safety guarantees while preserving high performance, grounded in rigorous theoretical foundations. Our framework includes: (1) an optimization component with chance constraints, (2) a safety classification model, (3) internal test data, (4) conservative testing procedures, (5) informative dataset quality measures, and (6) continuous approximate loss functions with gradient computation. Furthermore, to our knowledge, we mathematically establish the first scaling law in AI safety research, relating data quantity to safety-performance trade-offs. Experiments across reinforcement learning, natural language generation, and production planning validate our framework and demonstrate superior performance. Notably, in reinforcement learning, we achieve 3 collisions during 10M actions, compared with 1,000-3,000 for PPO-Lag baselines at equivalent performance levels -- a safety level unattainable by previous AI methods. We believe our framework opens a new foundation for safe AI deployment across safety-critical domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。