arXiv:2606.16022cs.RO2026-06

提出新方法提升人形机器人安全控制的边界判断与安全裕度估计

$λ$-Reachability: Geometric-Horizon Safety Bellman Equations for Humanoid Safety

论文配图:$λ$-Reachability: Geometric-Horizon Safety Bellman Equations for Humanoid Safety
图 1 · 摘自论文原文
  • 用几何分布滚动步数和随机终止点估算安全值,替代传统单步更新
  • 在模拟与真实人形机器人上,安全集边界分类准确率显著优于基线
  • 可通过参数调控安全分析范围,适合高维复杂系统的实时安全学习

我们提出 $λ$-Reachability,一种可扩展的高维机器人系统哈密顿-雅可比安全分析方法。不同于依赖固定单步贝尔曼更新的已有折扣方法,$λ$-Reachability 采用随机多步安全值估计器,结合几何分布的展开时长和随机吸收终止点。概念上类似于 TD($λ$),通过可解释的时长控制参数,在局部自洽更新与长时程轨迹最大安全性目标间插值。与 TD($λ$) 不同,$λ$-Reachability 中的终端安全值仅以概率 $δ$ 被使用;当 $δ<1$ 时,更新诱导收缩映射,支持时序差分学习;当 $λ\to 1$ 时,估计器恢复无折扣可达性目标。我们将该方法应用于受平衡与碰撞避让约束的高维安全学习问题,涵盖模拟与真实人形机器人。实验表明,相比单步时序差分基线,$λ$-Reachability 在安全集边界分类和安全裕度估计方面均有显著提升。

原文摘要 · Abstract (English)

We introduce $λ$-Reachability, a scalable approach to Hamilton--Jacobi safety analysis for high-dimensional robotic systems. Unlike prior discounted formulations that rely on fixed one-step Bellman updates, $λ$-Reachability employs a stochastic multi-step estimator of the safety value, using a geometrically distributed rollout horizon together with a randomly absorbed terminal. Conceptually analogous to TD($λ$), $λ$-Reachability interpolates between local self-consistency updates and long-horizon max-over-trajectory safety targets via an interpretable horizon-control parameter. Unlike TD($λ$), where the terminal value is always incorporated in learning targets, the terminal safety value in $λ$-Reachability is only used at a probability controlled by parameter $δ$. We formally show that for $δ<1$, the update induces a contraction mapping that allows temporal-difference learning; as $λ\to 1$, the estimator recovers the undiscounted reachability objective. We apply $λ$-Reachability to high-dimensional safety learning problems with both simulated and real humanoid robots under balance and collision avoidance constraints. Experimental results demonstrate that $λ$-Reachability significantly improves both safe-set boundary classification and safety margin estimation compared to single-step temporal-difference baselines.

机器人安全强化学习可达性分析人形机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。