arXiv:2502.00835cs.ROcs.LG2025-02被引 1

让机器人通过因果影响机制高效学会推物,无需大量试错。

CAIMAN: Causal Action Influence Detection for Sample-efficient Loco-manipulation

  • 用因果动作影响做内在激励,驱动机器人自主学习推物。
  • 在仿真中仅用少量样本就掌握推物技能,真实机器人可直接部署。
  • 适合研究样本高效强化学习与复杂环境下的机器人操控。

使腿式机器人执行非抓取式运动操作对提升其灵活性至关重要。学习如全身推物等行为通常需要复杂的规划策略或大量特定任务的奖励设计,尤其在非结构化环境中。本文提出CAIMAN,一种实用的强化学习框架,促使智能体获得对环境内其他实体的控制力。该方法利用因果动作影响作为内在动机目标,使腿式机器人在稀疏任务奖励下仍能高效习得物体推移技能。采用分层控制策略,结合低层运动模块与高层生成任务相关速度指令的策略,后者以最大化内在奖励为目标进行训练。为估计因果动作影响,通过整合运动学先验与训练期间收集的数据来学习环境动态。实验表明,CAIMAN在仿真中展现出卓越的样本效率和对多种场景的适应性,并成功迁移至真实系统而无需额外微调。视频演示见:https://www.youtube.com/watch?v=dNyvT04Cqaw。

原文摘要 · Abstract (English)

Enabling legged robots to perform non-prehensile loco-manipulation is crucial for enhancing their versatility. Learning behaviors such as whole-body object pushing often requires sophisticated planning strategies or extensive task-specific reward shaping, especially in unstructured environments. In this work, we present CAIMAN, a practical reinforcement learning framework that encourages the agent to gain control over other entities in the environment. CAIMAN leverages causal action influence as an intrinsic motivation objective, allowing legged robots to efficiently acquire object pushing skills even under sparse task rewards. We employ a hierarchical control strategy, combining a low-level locomotion module with a high-level policy that generates task-relevant velocity commands and is trained to maximize the intrinsic reward. To estimate causal action influence, we learn the dynamics of the environment by integrating a kinematic prior with data collected during training. We empirically demonstrate CAIMAN's superior sample efficiency and adaptability to diverse scenarios in simulation, as well as its successful transfer to real-world systems without further fine-tuning. A video demo is available at https://www.youtube.com/watch?v=dNyvT04Cqaw.

强化学习机器人操控样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。