arXiv:2508.03526cs.RO2025-08被引 5

让多机器人协作搬大件物体更智能,支持不同任务和机器人数目。

CollaBot: Vision-Language Guided Simultaneous Collaborative Manipulation

  • 用视觉语言模型定位目标,分步生成抓取姿势并协调多机动作
  • 在多种场景下实现72%成功率,优于传统行为克隆方法
  • 适合工厂、家庭等复杂环境中多机器人协同搬运任务

机器人与物理世界的交互是核心目标。传统操作研究多聚焦单机器人与小物体,但工厂与家庭环境常需搬运如桌子等大型物件,要求多机器人协同。现有研究缺乏可泛化处理多样化物体、任务与机器人数量的框架。本文提出CollaBot,一种通用的同步协作操作框架。首先利用SEEM进行场景分割与目标提取;其次设计协作抓取框架,将任务分解为局部抓取姿态生成与全局协调;最后构建两阶段规划模块,生成无碰撞执行轨迹。在不同物体、任务及机器人数量下的实验显示,该框架成功率高达72%,显著优于基于行为克隆的方法,验证了其在复杂多机器人协作任务中的优势。真实世界实验进一步证明了该方法在实际应用中的可行性。

原文摘要 · Abstract (English)

One central goal of robotics is to enable robots to interact with the physical world. Traditional manipulation studies primarily focus on single robots and relatively small objects. However, factory and domestic environments often require large-object manipulation, such as moving tables, where multiple robots must work collaboratively. Existing studies still lack a generalizable framework that can handle diverse objects, tasks, and robot team sizes. In this work, we propose CollaBot, a generalist framework for simultaneous collaborative manipulation. First, we use SEEM for scene segmentation and target-object extraction. Then, we propose a collaborative grasping framework that decomposes the task into local grasp pose generation and global coordination. Finally, we design a two-stage planning module to generate collision-free trajectories for task execution. Experimental results across different settings with varying objects, tasks, and numbers of robots indicate that our framework achieves a 72% success rate. This marks a substantial improvement over behavior cloning-based methods, validating the advantages of the proposed framework in complex multi-robot cooperative tasks. Real-world experiments further demonstrate the feasibility of our method in practical applications.

多机器人协作操作视觉语言抓取规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。