arXiv:2603.09971cs.RO2026-03被引 9

无需训练数据,用视觉和语言直接指挥机器人完成操作任务。

TiPToP: A Modular Open-Vocabulary Robot Manipulation System That Plans

  • 模块化设计整合感知、规划与执行,通过预训练模型直接处理图像和自然语言。
  • 在真实机器人上平均成功率更高,完成时间更快,优于微调过的先进系统。
  • 可快速部署且便于故障定位,适合研究模块化机器人的开发者使用。

我们提出TiPToP,一个模块化的开放词汇机器人操作系统,将预训练基础模型与GPU加速的任务和运动规划器结合,直接从RGB图像和自然语言中解决任务。TiPToP由感知、规划和执行模块组成,无需机器人训练数据,可在标准DROID配置下一小时内部署,并以极小成本适配新机器人形态。我们在两个真实DROID设置(其中一个由外部团队操作)和仿真环境中评估,结果表明其平均成功率高于$π_{0.5}\text{-DROID}$(该系统基于350小时演示微调),且平均完成时间更短。在MolmoSpaces基准测试中,未在分布内数据上训练的模型中,TiPToP在抓取和抓取-放置任务上排名第一。此外,其模块化结构支持失败溯源,有助于精准改进。项目代码与网站已开源,作为可复现基线,推动模块化操作系统的进一步研究。

原文摘要 · Abstract (English)

We present TiPToP, a modular manipulation system that integrates pretrained foundation models with a GPU-accelerated Task and Motion Planner to solve tasks directly from RGB images and natural language. TiPToP composes perception, planning, and execution modules and requires no robot training data. It can be deployed on a standard DROID setup in under an hour and adapted to new embodiments with minimal effort. We evaluate TiPToP against $π_{0.5}\text{-DROID}$, a state-of-the-art VLA fine-tuned on 350 hours of demonstrations, across two real-world DROID setups (one operated by an external team) and simulation, where TiPToP attains a higher average success rate and faster average completion time. We also evaluate on the MolmoSpaces benchmark, where TiPToP ranks first overall on pick and pick-and-place tasks among methods not trained on in-distribution data. We further show that TiPToP's modularity enables us to trace failures to specific components, revealing where to target improvements. We release TiPToP open-source to serve as a reproducible baseline and to enable further research on modular manipulation systems. Project website and code: https://tiptop-robot.github.io

机器人操作模块化系统视觉语言模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。