arXiv:2608.28090cs.RO2026-08

让机器人坐椅子上实现全向移动,不依赖触觉反馈也能精准跟随时速指令。

Stay Seated: Learning Omnidirectional Humanoid Locomotion on a Passive Mobile Chair with Casters

论文配图:Stay Seated: Learning Omnidirectional Humanoid Locomotion on a Passive Mobile Chair with Casters
图 1 · 摘自论文原文
  • 用被动轮椅模型模拟坐姿运动,仅靠本体感知和速度指令训练策略。
  • 在随机指令测试中,95%以上轨迹能完成20秒全向运动,性能超越站立模式。
  • 无需接触传感,零样本迁移到真实机器人,适合人机协同场景研究。

采用准直接驱动关节的人形机器人在站立时持续产生关节力矩,而坐着的人类则将体重支撑交给椅子。作为坐姿运动操作的初步探索,本文研究了在被动移动轮椅上的全向坐姿行走,要求髋部与座椅无固定接触,并间歇性地通过脚与地面的推进来驱动机器人-轮椅系统。我们扩展了标准的站立速度跟踪环境,引入被动轮椅模型、坐姿奖励机制、仅由评判器观测的轮椅状态以及任务定制化的接触设置。策略在没有运动模仿奖励的情况下学习;其执行器仅使用本体感知和速度指令,不依赖接触感知或轮椅状态信息。在随机指令评估中,策略在近20秒的轨迹中成功追踪全向指令,最佳坐姿策略在速度追踪表现上优于站立策略。通过对四个训练种子进行$2^3$全因子实验分析对称性正则化(SY)、足滑正则化(FS)和指令课程(CC)的影响,结果显示FS降低了能量消耗(CoT),但增加了追踪误差;部分仅含FS的策略收敛至静止局部最优解。结合FS与任一其他方法(SY或CC)可避免该问题,且无需重新调参;而SY在纵向运动中提升了双侧腿部对称性。方向解析分析显示,能量消耗顺序为:后向 < 侧向 ≪ 前向,其中后向和侧向运动中存在植腿伸展,前向运动中踝接触后伴随膝屈曲。所学策略实现了零样本从仿真到现实的迁移,在Unitree G1机器人上生成了全向坐姿运动。

原文摘要 · Abstract (English)

Humanoid robots with quasi-direct-drive actuators continuously generate joint torque while standing, whereas seated humans delegate weight support to chairs during desk work. As a first step toward seated loco-manipulation, we study omnidirectional seated locomotion on a passive mobile chair, requiring unfixed pelvis-seat contact and intermittent foot-floor propulsion of the robot-chair system. We extend a standard standing velocity-tracking environment with a passive-chair model, seated-state rewards, critic-only chair observations, and task-tailored contact settings. The policy is learned without motion-imitation rewards; its actor uses only proprioception and velocity commands, without contact sensing or chair states. In random-command evaluation, the policies tracked omnidirectional commands through nearly all 20-s rollouts, and the best seated policies could outperform the Standing policy in velocity tracking. Across four training seeds, a $2^3$ full-factorial comparison of symmetry regularization (SY), foot-slip regularization (FS), and command curriculum (CC) showed that FS reduced CoT but increased tracking error and that some FS-only policies converged to stationary local optima. Combining FS with either SY or CC avoided this failure without retuning FS, while SY improved bilateral leg symmetry during longitudinal motion. Direction-resolved analysis showed CoT ordered backward $<$ lateral $\ll$ forward, with planted-leg extension in backward and lateral motion and knee flexion following heel contact in forward motion. The learned policy achieved zero-shot sim-to-real transfer to a Unitree G1 and generated omnidirectional seated locomotion.

人形机器人坐姿运动强化学习零样本迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。