用视觉语言模型自动设计机器人工具并规划操作,提升物理任务解决能力。
VLMgineer: Vision Language Models as Robotic Toolsmiths
- 结合视觉语言模型与进化搜索,协同生成工具和使用策略。
- 在多样日常操作任务中,工具设计成功率超人类手工工具和单纯模型生成方案。
- 适合对机器人自主创造、具身智能感兴趣的学者和工程师。
工具的设计与使用体现了理解并操控物理世界所需的创造力、规划与远见,常被视为衡量智能的重要指标。当前机器人研究多聚焦于优化控制策略,而发明更聪明的工具则提供了另一种物理智能形式:将问题求解的负担转移至工具本身。鉴于现代基础模型在常识、推理和创造力方面的强大能力,我们探究这些模型能否为自动设计与高效使用工具提供有效先验知识。本文提出 VLMgineer 框架,利用视觉语言模型(VLMs)的代码生成能力,结合进化搜索,迭代地协同设计物理工具及其操作动作策略以完成特定任务。我们在一个涵盖多种日常操作场景的新基准上评估该框架,结果显示,VLMgineer 在各类任务中均能持续发现更有效且更具创新性的工具与策略,将复杂机器人问题转化为简单执行任务。其表现优于基于人类指令生成的工具设计以及现有手工制作工具。为促进自动化工具发明研究,我们将公开本研究的基准数据集与代码。
原文摘要 · Abstract (English)
Tool design and use reflect the ability to understand and manipulate the physical world through creativity, planning, and foresight. As such, these capabilities are often regarded as measurable indicators of intelligence across biological species. While much of today's research on robotic intelligence focuses on generating better controllers, inventing smarter tools offers a complementary form of physical intelligence: shifting the onus of problem-solving onto the tool's design. Given the vast and impressive common-sense, reasoning, and creative capabilities of today's foundation models, we investigate whether these models can provide useful priors to automatically design and effectively wield such tools? We present VLMgineer, a framework that harnesses the code generation abilities of vision language models (VLMs) together with evolutionary search to iteratively co-design physical tools and the action plans that operate them to perform a task. We evaluate VLMgineer on a diverse new benchmark of everyday manipulation scenarios that demand creative tool design and use. Across this suite, VLMgineer consistently discovers tools and policies that solve tasks more effectively and innovatively, transforming challenging robotics problems into straightforward executions. It also outperforms VLM-generated designs from human specifications and existing human-crafted tools for everyday tasks. To facilitate future research on automated tool invention, we will release our benchmark and code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。