arXiv:2503.23271cs.ROcs.AI2025-03ICRA被引 11

用状态扩散与逆动力学模型,让机器人双手协同操作更自然精准。

Learning Coordinated Bimanual Manipulation Policies using State Diffusion and Inverse Dynamics Models

  • 分离物体状态变化与机器人动作模型,提升双手协作能力
  • 在多种仿真与真实场景中表现优于现有方法,尤其擅长处理多目标和变形体
  • 适合研究双臂机器人控制、具身智能与模仿学习的开发者

执行如洗衣这类任务时,人类能自然协调双手操控衣物并预判动作对衣物状态的影响。然而,机器人实现这种协调仍面临建模物体运动、预测未来状态和生成精确双臂动作的挑战。本文通过将人类操作策略的预测性融入机器人模仿学习,提出新方法:将任务相关的状态转移与代理特异的逆动力学模型解耦,以实现有效的双臂协作。利用演示数据集,训练一个扩散模型来根据历史观测预测未来状态,模拟场景演化;再通过逆动力学模型计算达成预测状态所需的机器人动作。关键发现是,建模物体运动有助于学习双臂协调控制策略。在包含多模态目标配置、双臂操作、可变形物体及多物体设置的多样化仿真与真实世界任务中,该框架始终优于当前最先进的状态到动作映射策略。方法展现出强大处理多模态目标配置与动作分布的能力,能在不同控制模式下保持稳定,并生成超越示范数据集范围的行为组合。

原文摘要 · Abstract (English)

When performing tasks like laundry, humans naturally coordinate both hands to manipulate objects and anticipate how their actions will change the state of the clothes. However, achieving such coordination in robotics remains challenging due to the need to model object movement, predict future states, and generate precise bimanual actions. In this work, we address these challenges by infusing the predictive nature of human manipulation strategies into robot imitation learning. Specifically, we disentangle task-related state transitions from agent-specific inverse dynamics modeling to enable effective bimanual coordination. Using a demonstration dataset, we train a diffusion model to predict future states given historical observations, envisioning how the scene evolves. Then, we use an inverse dynamics model to compute robot actions that achieve the predicted states. Our key insight is that modeling object movement can help learning policies for bimanual coordination manipulation tasks. Evaluating our framework across diverse simulation and real-world manipulation setups, including multimodal goal configurations, bimanual manipulation, deformable objects, and multi-object setups, we find that it consistently outperforms state-of-the-art state-to-action mapping policies. Our method demonstrates a remarkable capacity to navigate multimodal goal configurations and action distributions, maintain stability across different control modes, and synthesize a broader range of behaviors than those present in the demonstration dataset.

双臂控制模仿学习扩散模型机器人操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。