用SDP方法训练出高维输入下可证明安全的神经网络分类器。
Training Safe Neural Networks with Global SDP Bounds
- 基于ADMM的训练方案,结合半定规划验证安全边界。
- 在40维输入下实现可证明的完美召回率。
- 适合需要形式化安全保障的高维系统,如安全强化学习。
本文提出一种基于半定规划(SDP)的神经网络训练新方法,可为高维输入区域提供形式化安全保证。该方法突破了现有技术仅关注对抗鲁棒性边界的局限,专注于大规模高维输入的安全验证。我们设计了一种基于交替方向乘子法(ADMM)的训练方案,在对抗球体数据集(Adversarial Spheres)上实现了准确的神经网络分类器,在输入维度高达 d=40 时仍能保证完美召回率。本工作推动了高维系统中可靠神经网络验证方法的发展,潜在应用于安全强化学习策略设计。
原文摘要 · Abstract (English)
This paper presents a novel approach to training neural networks with formal safety guarantees using semidefinite programming (SDP) for verification. Our method focuses on verifying safety over large, high-dimensional input regions, addressing limitations of existing techniques that focus on adversarial robustness bounds. We introduce an ADMM-based training scheme for an accurate neural network classifier on the Adversarial Spheres dataset, achieving provably perfect recall with input dimensions up to $d=40$. This work advances the development of reliable neural network verification methods for high-dimensional systems, with potential applications in safe RL policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。