通过生成多样泊位示范,提升机器人移动操作的视角泛化能力。
DockAnywhere: Data-Efficient Visuomotor Policy Learning for Mobile Manipulation via Novel Demonstration Generation

- 将基础运动与抓取技能解耦,从单个示范生成多泊位轨迹。
- 在ManiSkill和真实平台测试中,新泊位成功率显著提升。
- 适合需要高泛化性的移动操作任务,如家庭或工厂机器人。
移动操作是机器人在家庭、工厂等广阔环境中交互的基础能力。现有方法多采用两阶段范式:先导航至固定泊位,再执行基于视觉的固定基操作。然而,实际应用中因泊位位置变化导致视角泛化困难。为此,本文提出低成本示范生成框架DockAnywhere,通过将单一示范提升至多种可行泊位配置,增强在泊位变化下的视角泛化能力。具体地,该方法将依赖泊位的基座运动与跨视角保持不变的接触密集型操作技能解耦;在可行性约束下采样可行泊位,并通过结构保持增强生成对应轨迹;利用点云表示机器人与物体,在3D空间中进行点级空间编辑,确保不同视角下观测与动作的一致性。在ManiSkill及真实平台上的大量实验表明,DockAnywhere显著提升策略成功率达数倍,且能有效泛化至训练中未见的泊位视角,大幅增强移动操作策略在真实部署中的泛化能力。
原文摘要 · Abstract (English)
Mobile manipulation is a fundamental capability that enables robots to interact in expansive environments such as homes and factories. Most existing approaches follow a two-stage paradigm, where the robot first navigates to a docking point and then performs fixed-base manipulation using powerful visuomotor policies. However, real-world mobile manipulation often suffers from the view generalization problem due to shifts of docking points. To address this issue, we propose a novel low-cost demonstration generation framework named DockAnywhere, which improves viewpoint generalization under docking variability by lifting a single demonstration to diverse feasible docking configurations. Specifically, DockAnywhere lifts a trajectory to any feasible docking points by decoupling docking-dependent base motions from contact-rich manipulation skills that remain invariant across viewpoints. Feasible docking proposals are sampled under feasibility constraints, and corresponding trajectories are generated via structure-preserving augmentation. Visual observations are synthesized in 3D space by representing the robot and objects as point clouds and applying point-level spatial editing to ensure the consistency of observation and action across viewpoints. Extensive experiments on ManiSkill and real-world platforms demonstrate that DockAnywhere substantially improves policy success rates and easily generalizes to novel viewpoints from unseen docking points during training, significantly enhancing the generalization capability of mobile manipulation policy in real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。