让机器人自动生成能用的工具,解决复杂操作难题。
RobotSmith: Generative Robotic Tool Design for Acquisition of Complex Manipulation Skills
- 用视觉语言模型和物理仿真协同设计工具,迭代优化形状与使用方式。
- 在多种任务中平均成功率50.0%,远超3D生成(21.4%)和工具检索(11.1%)。
- 适合需要自主工具设计的机器人应用场景,如工业装配、家庭服务。
赋予机器人自主设计工具的能力,是其完成复杂操作任务的关键。现有生成框架虽可自动合成3D场景与奖励函数,但尚未解决工具使用场景的设计问题。直接调用人类设计的工具往往不适用,因许多工具(如擀面杖)难以被机械臂操作。现有方法或依赖固定模板、参数调整有限,或使用通用3D生成技术,未针对工具生成优化。为此,我们提出RobotSmith,一个自动化流水线,融合视觉语言模型(VLMs)中的隐式物理知识与物理仿真的高精度物理模拟,实现工具设计与使用。系统通过协作式VLM代理迭代提出工具构型,生成低层机器人轨迹,并联合优化工具几何形状与使用策略以提升任务表现。我们在包含刚体、柔体及流体对象的多种操作任务上评估该方法,实验表明其在任务成功率与整体性能上均显著优于强基线。尤其在平均成功率上达50.0%,大幅超越3D生成(21.4%)与工具检索(11.1%)。最后,我们在真实环境中部署系统,验证生成工具及其使用方案可有效迁移到物理执行,证明了本方法的实用性与泛化能力。
原文摘要 · Abstract (English)
Endowing robots with tool design abilities is critical for enabling them to solve complex manipulation tasks that would otherwise be intractable. While recent generative frameworks can automatically synthesize task settings, such as 3D scenes and reward functions, they have not yet addressed the challenge of tool-use scenarios. Simply retrieving human-designed tools might not be ideal since many tools (e.g., a rolling pin) are difficult for robotic manipulators to handle. Furthermore, existing tool design approaches either rely on predefined templates with limited parameter tuning or apply generic 3D generation methods that are not optimized for tool creation. To address these limitations, we propose RobotSmith, an automated pipeline that leverages the implicit physical knowledge embedded in vision-language models (VLMs) alongside the more accurate physics provided by physics simulations to design and use tools for robotic manipulation. Our system (1) iteratively proposes tool designs using collaborative VLM agents, (2) generates low-level robot trajectories for tool use, and (3) jointly optimizes tool geometry and usage for task performance. We evaluate our approach across a wide range of manipulation tasks involving rigid, deformable, and fluid objects. Experiments show that our method consistently outperforms strong baselines in terms of both task success rate and overall performance. Notably, our approach achieves a 50.0\% average success rate, significantly surpassing other baselines such as 3D generation (21.4%) and tool retrieval (11.1%). Finally, we deploy our system in real-world settings, demonstrating that the generated tools and their usage plans transfer effectively to physical execution, validating the practicality and generalization capabilities of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。