提出可并行的可微可达性框架,让神经网络控制在机器人中更安全可靠。
Parallel Differentiable Reachability for Learning and Planning with Certified Neural Dynamics and Controllers

- 用泰勒模型与线性传播构建可微分的可达集计算方法
- 在72维硬件实验中实现在线规划与认证覆盖
- 适合做安全强化学习和实时控制的工程师使用
神经网络动力学模型和控制策略在机器人领域表现优异,但在不确定性下提供可靠保障仍具挑战,尤其针对闭环神经网络系统。现有可达性工具虽能给出形式化上界,但通常不可微、过于保守或速度过慢,难以适配现代学习与在线规划流程。本文提出一种基于JAX的并行可微可达性框架,适用于连续与离散时间系统,支持解析与神经网络动力学及控制器。该框架结合泰勒模型流管构造与CROWN风格的线性边界传播,采用统一表示以保留仿射依赖关系,同时支持GPU批量计算与自动微分。基于此可达性基元,我们开发了(i)一种认证训练方法,促使动态模型与控制器具备良好的可达性特性;(ii)一种感知可达性的采样型模型预测控制方案,支持梯度优化改进。在非抓取操作与四旋翼任务中的实验,包括硬件测试与高达72维的高维评估,验证了该方法在保持有界不确定性下认证可达集上界的同时,实现了实用的在线规划能力。
原文摘要 · Abstract (English)
Neural network (NN) dynamics models and control policies achieve strong performance in robotics, but providing sound guarantees under uncertainty remains difficult, especially for closed-loop NN systems. Existing reachability tools provide formal over-approximations, yet are often non-differentiable, overly conservative, or too slow for modern learning and online planning pipelines. To address this, we present a parallelizable, differentiable reachability framework in JAX for continuous- and discrete-time systems with analytical and NN-based dynamics and controllers. Our framework combines Taylor-model flowpipe construction with CROWN-style linear bound propagation through a unified representation that preserves affine dependencies while supporting GPU-batched computation and automatic differentiation. Building on this reachability primitive, we develop (i) a certified training method that encourages reachability-friendly dynamics models and controllers, and (ii) a reachability-aware sampling-based MPC scheme with gradient-based refinement. Experiments on non-prehensile manipulation and quadrotor tasks, including hardware and higher-dimensional evaluations (up to 72D), demonstrate practical online planning while maintaining certified reachable-set over-approximations under bounded uncertainty.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。