arXiv:2608.20823cs.RO2026-08

用强化学习让机器人自然起身,无需示例动作。

Natural Sit-to-Stand Motion Synthesis For Humanoids via Guided Assistance Curricula and Staged Rewards

论文配图:Natural Sit-to-Stand Motion Synthesis For Humanoids via Guided Assistance Curricula and Staged Rewards
图 1 · 摘自论文原文
  • 分阶段训练:先用辅助力和矮椅帮助探索,逐步提升难度。
  • 在8种不同椅高下成功率超97%,能从深坐姿势平稳站起。
  • 结合生物力学设计奖励机制,动作流畅省力,适合仿人机器人。

类人机器人从坐姿起身时有无限种平衡方式,使该任务成为复杂的控制难题。本文通过强化学习从零开始合成自然的类人机器人起身动作,无需示范或参考轨迹。单一近端策略优化(Proximal Policy Optimization)策略在三个互补组件驱动下学习到平滑、类人的起身动作:(i) 联合使用力辅助与椅高递进课程,垂直骨盆辅助力初期介入并随训练衰减,椅高逐步解锁,确保模型在每种高度掌握可行轨迹后再进入更难情形,避免过早分布偏移导致泛化崩溃;(ii) 通过大量逆运动学生成的初始与目标姿态随机采样,覆盖超过八种椅高,提升运动鲁棒性;(iii) 奖励函数受生物力学与最优控制研究启发,调控离座瞬间角动量,并通过质心压力吸引函数实现支撑区域平稳过渡,确保低能耗执行。在无外力确定性评估器上,该策略在八种椅高下成功率达97%以上,具备跨椅高的平滑泛化能力,且可处理远比现有方法更深的坐姿状态。

原文摘要 · Abstract (English)

A humanoid has infinitely many ways to stand up from sitting while maintaining balance, making sit-to-stand (STS) a challenging control problem. We synthesise natural humanoid STS motion from scratch using reinforcement learning, without demonstrations or reference trajectories. A single Proximal Policy Optimisation policy learns smooth, human-like rising driven by three complementary components. (i) A coupled force/chair-height curriculum is used. A vertical pelvis-assist force aids early trajectory exploration and decays over training. Taller chairs are unlocked with decaying assisting force. This ensures that the policy masters a viable STS trajectory at each chair height before being exposed to harder ones, avoiding the premature distribution shift that otherwise collapses generalisation. (ii) Motion robustness is achieved by randomly sampling from a large number of inverse kinematics-generated initial and target poses spanning over eight chair heights. (iii) A set of rewards is defined inspired from biomechanics and optimal control studies. They shape the robot's angular momentum for seat-off, and enable support-region transition via centre of pressure attraction function to ensure smooth low-effort actuation. On a deterministic force-free evaluator, the policy attains more than 97% balanced-standing success across eight chair heights. The policy generalises smooth motion across chair heights and enables the robot to rise from substantially deep-seated postures as compared to the state of the art.

强化学习机器人控制动作合成类人机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。