arXiv:2506.12095cs.RO2025-06被引 1

提出双感知机制,让机器人在复杂地形中更安全高效地学习行走。

DoublyAware: Dual Planning and Policy Awareness for Temporal Difference Learning in Humanoid Locomotion

  • 分离规划与策略不确定性,用置信区间筛选可靠轨迹
  • 在26自由度人形机器人上实现更快收敛和更高运动可行性
  • 适合关注强化学习稳定性与数据效率的研究者

实现人形机器人稳健运动学习是基于模型的强化学习(MBRL)的核心挑战,环境中的随机性(即偶然不确定性)在高维动作空间与复杂接触动力学下会被放大,并与模型的信念不确定性(认知不确定性)交织,影响探索效率与学习稳定性。本文提出DoublyAware,一种面向时序差分模型预测控制(TD-MPC)的不确定性感知扩展方法,将不确定性分解为可解释的两个独立分量:规划不确定性和策略不确定性。为应对规划不确定性,DoublyAware采用符合性预测(conformal prediction)技术,通过量化校准的风险边界筛选候选轨迹,确保统计一致性与对随机动力学的鲁棒性。同时,利用策略滚动采样作为结构化先验,结合组相对策略约束(GRPC)优化器,在潜在动作空间中施加基于群体的自适应信任区域。该协同设计使智能体能优先选择高置信度、高回报行为,同时在不确定性下保持有效且精准的探索。在包含Unitree 26-DoF H1-2人形机器人的HumanoidBench运动基准测试中,DoublyAware展现出更高的样本效率、更快的收敛速度和更强的运动可行性,验证了结构化不确定性建模对数据高效与可靠决策的关键作用。

原文摘要 · Abstract (English)

Achieving robust robot learning for humanoid locomotion is a fundamental challenge in model-based reinforcement learning (MBRL), where environmental stochasticity and randomness can hinder efficient exploration and learning stability. The environmental, so-called aleatoric, uncertainty can be amplified in high-dimensional action spaces with complex contact dynamics, and further entangled with epistemic uncertainty in the models during learning phases. In this work, we propose DoublyAware, an uncertainty-aware extension of Temporal Difference Model Predictive Control (TD-MPC) that explicitly decomposes uncertainty into two disjoint interpretable components, i.e., planning and policy uncertainties. To handle the planning uncertainty, DoublyAware employs conformal prediction to filter candidate trajectories using quantile-calibrated risk bounds, ensuring statistical consistency and robustness against stochastic dynamics. Meanwhile, policy rollouts are leveraged as structured informative priors to support the learning phase with Group-Relative Policy Constraint (GRPC) optimizers that impose a group-based adaptive trust-region in the latent action space. This principled combination enables the robot agent to prioritize high-confidence, high-reward behavior while maintaining effective, targeted exploration under uncertainty. Evaluated on the HumanoidBench locomotion suite with the Unitree 26-DoF H1-2 humanoid, DoublyAware demonstrates improved sample efficiency, accelerated convergence, and enhanced motion feasibility compared to RL baselines. Our simulation results emphasize the significance of structured uncertainty modeling for data-efficient and reliable decision-making in TD-MPC-based humanoid locomotion learning.

人形机器人强化学习不确定性建模运动控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。