arXiv:2606.00252cs.ROcs.LG2026-06

用模仿学习+高效强化学习,让人形机器人精准操控晃动的吊载物。

HOIST: Humanoid Optimization with Imitation and Sample-efficient Tuning for Manipulating Suspended Loads

论文配图:HOIST: Humanoid Optimization with Imitation and Sample-efficient Tuning for Manipulating Suspended Loads
图 1 · 摘自论文原文
  • 先用虚拟现实演示微调视觉语言动作策略,再通过批量强化学习优化定位精度。
  • 实机与仿真测试显示,定位误差减少19.9厘米,角度误差降低3.56度。
  • 适合需要高精度吊装作业的人形机器人应用,如工厂搬运或灾难救援。

用人形机器人操控悬吊负载极具挑战性,因机器人仅能通过全身运动和间歇接触影响欠驱动、振荡的负载。模仿学习可提供安全初始行为,但无法直接优化最终放置位置;而从零开始的强化学习在真实人形机器人上既不安全又样本效率低。本文提出HOIST——基于模仿学习与高效采样强化学习的人形机器人悬吊负载操作方法。首先,利用虚拟现实遥操作示范微调高层视觉-语言-动作(VLA)策略,并通过全身控制器执行指令;随后,结合VLA回放与迭代批量强化学习,提升放置精度与停止行为。仿真与真实人形机器人实验表明,相比仅使用模仿学习及额外示范基线,HOIST显著提升性能;相较于纯VLA回放,其平移放置误差减少19.9厘米,原始角度误差降低3.56度,验证了人形机器人在欠驱动物料搬运任务中的潜力。

原文摘要 · Abstract (English)

Manipulating suspended payloads with humanoid robots is challenging because the robot can only influence an underactuated, oscillatory load through whole-body motion and intermittent contact. Imitation learning provides safe initial behavior but does not directly optimize final placement, while reinforcement learning from scratch is unsafe and sample-inefficient on real humanoids. We present HOIST-Humanoid Optimized with Imitation and Sample-efficient Tuning for manipulating suspended loads. HOIST first finetunes a high-level vision-language-action (VLA) policy from virtual-reality (VR) teleoperation demonstrations and executes its commands through a whole-body controller. It then uses VLA rollouts and iterative batched RL to improve placement accuracy and stopping behavior. Experiments in simulation and on a real humanoid show that HOIST improves over imitation-only and additional-demonstration baselines; compared with pure VLA rollouts, HOIST reduces translational placement error by 19.9 cm and raw angular error by 3.56 degrees, demonstrating the potential of humanoids for underactuated material-handling tasks.

人形机器人强化学习操作控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。