用引导式强化学习实现软体机器人零样本全身操控,成功率达88%
Zero-shot Whole-Body Manipulation with a Large-Scale Soft Robotic Torso via Guided Reinforcement Learning
- 用简单运动基元引导强化学习,突破传统奖励设计瓶颈
- 在10公斤负载下实现88%成功率的零样本仿真到现实迁移
- 策略具自适应反应能力,可主动重抓与抗扰恢复
全身操作能令机器人通过非末端执行器部分与大型、重型或不规则物体交互,但其对软体机器人而言挑战巨大,因动力学与运动学不确定性高。本文构建可在单线程上以350倍真实速度运行的模拟环境(基于MuJoCo),并分析了速度与精度间的权衡。基于此框架,实现了在Baloo硬件平台上的零样本仿真到现实迁移,成功率高达88%。结果表明,用简单运动基元引导强化学习是成功关键,而标准奖励设计难以生成稳定有效的全身操控策略。所学策略不仅不机械模仿基元,还展现出有益的动态响应行为,如重新抓取和扰动恢复。对比开环基线发现,该策略在扰动下亦可能产生激进过纠正行为。据我们所知,这是首个在大型平台(10公斤负载)上使用双连续体软臂实现六自由度强力全身操控且具备零样本迁移的成功案例。
原文摘要 · Abstract (English)
Whole-body manipulation is a powerful yet underexplored approach that enables robots to interact with large, heavy, or awkward objects using more than just their end-effectors. Soft robots, with their inherent passive compliance, are particularly well-suited for such contact-rich manipulation tasks, but their uncertainties in kinematics and dynamics pose significant challenges for simulation and control. In this work, we address this challenge with a simulation that can run up to 350x real time on a single thread in MuJoCo and provide a detailed analysis of the critical tradeoffs between speed and accuracy for this simulation. Using this framework, we demonstrate a successful zero-shot sim-to-real transfer of a learned whole-body manipulation policy, achieving an 88% success rate on the Baloo hardware platform. We show that guiding RL with a simple motion primitive is critical to this success where standard reward shaping methods struggled to produce a stable and successful policy for whole-body manipulation. Furthermore, our analysis reveals that the learned policy does not simply mimic the motion primitive. It exhibits beneficial reactive behavior, such as re-grasping and perturbation recovery. We analyze and contrast this learned policy against an open-loop baseline to show that the policy can also exhibit aggressive over-corrections under perturbation. To our knowledge, this is the first demonstration of forceful, six-DoF whole-body manipulation using two continuum soft arms on a large-scale platform (10 kg payloads), with zero-shot policy transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。