arXiv:2509.18648cs.ROcs.AI2025-09NeurIPS被引 8

用悲观域随机化让仿真训练的机器人安全落地。

SPiDR: A Simple Approach for Zero-Shot Safety in Sim-to-Real Transfer

  • 在域随机化中引入悲观安全约束,应对仿真到现实的偏差。
  • 实测证明在真实机器人上仍能保持高安全性和强性能。
  • 兼容现有训练流程,适合需安全落地的强化学习应用。

在真实世界部署强化学习面临挑战,因仿真中训练的策略必须应对不可避免的仿真到现实差距。稳健的安全强化学习方法虽有理论保障但难扩展,而域随机化更实用却易引发不安全行为。本文提出SPiDR(Sim-to-real via Pessimistic Domain Randomization),一种可扩展且具备理论保障的安全仿真到现实迁移算法。SPiDR通过域随机化将仿真到现实的不确定性融入安全约束,兼具灵活性与兼容性。在多个仿真-仿真基准及两个不同的真实机器人平台上进行大量实验,结果表明SPiDR在存在仿真到现实差距时仍能有效保证安全性,同时维持优异性能。

原文摘要 · Abstract (English)

Deploying reinforcement learning (RL) safely in the real world is challenging, as policies trained in simulators must face the inevitable sim-to-real gap. Robust safe RL techniques are provably safe, however difficult to scale, while domain randomization is more practical yet prone to unsafe behaviors. We address this gap by proposing SPiDR, short for Sim-to-real via Pessimistic Domain Randomization -- a scalable algorithm with provable guarantees for safe sim-to-real transfer. SPiDR uses domain randomization to incorporate the uncertainty about the sim-to-real gap into the safety constraints, making it versatile and highly compatible with existing training pipelines. Through extensive experiments on sim-to-sim benchmarks and two distinct real-world robotic platforms, we demonstrate that SPiDR effectively ensures safety despite the sim-to-real gap while maintaining strong performance.

强化学习安全控制仿真迁移机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。