arXiv:2409.01083cs.ROcs.AI2024-09被引 78

用流匹配统一物体可操作性与机械臂动作生成,提升家务场景机器人效率。

Affordance-based Robot Manipulation with Flow Matching

  • 通过可学习提示微调冻结视觉模型,高效适应多任务可操作性理解。
  • 采用监督流匹配学习动作轨迹,在多个基准上优于行为克隆方法。
  • 适合需要快速部署、多任务泛化的家庭服务机器人研发人员。

我们提出一种辅助机器人操作框架,解决两大挑战:一是高效将大规模模型适配到下游场景可操作性理解任务,尤其在涉及人类的日常场景中收集多任务数据成本高昂;二是通过视觉可操作性模型有效学习机器人动作轨迹。针对第一点,我们采用参数高效的提示微调方法,在冻结的视觉模型前添加可学习文本提示,以预测多任务场景下的可操作性。随后,我们提出一种基于监督流匹配的方法,以可操作性为引导学习机器人动作轨迹。流匹配将机器人视觉-运动策略建模为从随机目标点到期望动作轨迹的条件演化过程。最后,我们构建了一个包含10项日常生活任务的真实世界数据集来测试该框架。大量评估表明,所提出的提示微调方法在不同数据规模下表现优异,甚至超越部分微调方案,同时保持参数高效性;使用流匹配学习多任务动作轨迹在多个机器人操作基准上均优于某些行为克隆方法,具有更稳定的训练与评估过程、显著更快的推理速度,且与扩散策略相比具备相当的泛化能力,多数情况下表现略优。该框架无缝融合了可操作性学习与动作生成,实现统一建模。

原文摘要 · Abstract (English)

We present a framework for assistive robot manipulation, which focuses on two fundamental challenges: first, efficiently adapting large-scale models to downstream scene affordance understanding tasks, especially in daily living scenarios where gathering multi-task data involving humans requires strenuous effort; second, effectively learning robot action trajectories by grounding the visual affordance model. We tackle the first challenge by employing a parameter-efficient prompt tuning method that prepends learnable text prompts to the frozen vision model to predict manipulation affordances in multi-task scenarios. Then we propose to learn robot action trajectories guided by affordances in a supervised flow matching method. Flow matching represents a robot visuomotor policy as a conditional process of flowing random waypoints to desired robot action trajectories. Finally, we introduce a real-world dataset with 10 tasks across Activities of Daily Living to test our framework. Our extensive evaluation highlights that the proposed prompt tuning method for learning manipulation affordance achieves competitive performance and even outperforms some other finetuning protocols across data scales, while satisfying parameter efficiency. Learning multi-task robot action trajectories with flow matching leads to consistently favorable results in several robot manipulation benchmarks than some alternative behavior cloning methods. This includes more stable training and evaluation, and noticeably faster inference, while maintaining comparable generalization performance to diffusion policy, where flow matching performs marginally better in most cases. Our framework seamlessly unifies affordance learning and action generation with flow matching for robot manipulation.

机器人操作流匹配可操作性提示微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。