机器人能自动生成工具并学习使用,实现自主完成人类指令的任务。
Evolution 6.0: Robot Evolution through Generative Design
- 用视觉语言模型和3D生成模型让机器人自动设计专用工具。
- 工具生成成功率90%,动作生成在物理与视觉泛化上达83.5%。
- 适合研究智能机器人自主演化与具身智能的开发者参考。
我们提出一种新概念——进化6.0,即由生成式AI驱动的机器人进化。当机器人无法完成人类提出的任务时,它会自主设计所需工具,并学习如何使用以达成目标。该系统基于视觉-语言模型(VLM)、视觉-语言-动作模型(VLA)和文本到3D生成模型,包含两个核心模块:工具生成模块,从视觉与文本数据中制造任务专用工具;动作生成模块,将自然语言指令转化为机器人动作。系统融合QwenVLM进行环境理解、OpenVLA执行任务、Llama-Mesh生成3D工具。评估结果显示,工具生成成功率达90%,推理时间仅10秒;动作生成在物理与视觉泛化上达到83.5%,运动泛化70%,语义泛化37%。未来将聚焦双臂操作、任务能力扩展与环境理解增强,提升真实场景适应性。
原文摘要 · Abstract (English)
We propose a new concept, Evolution 6.0, which represents the evolution of robotics driven by Generative AI. When a robot lacks the necessary tools to accomplish a task requested by a human, it autonomously designs the required instruments and learns how to use them to achieve the goal. Evolution 6.0 is an autonomous robotic system powered by Vision-Language Models (VLMs), Vision-Language Action (VLA) models, and Text-to-3D generative models for tool design and task execution. The system comprises two key modules: the Tool Generation Module, which fabricates task-specific tools from visual and textual data, and the Action Generation Module, which converts natural language instructions into robotic actions. It integrates QwenVLM for environmental understanding, OpenVLA for task execution, and Llama-Mesh for 3D tool generation. Evaluation results demonstrate a 90% success rate for tool generation with a 10-second inference time, and action generation achieving 83.5% in physical and visual generalization, 70% in motion generalization, and 37% in semantic generalization. Future improvements will focus on bimanual manipulation, expanded task capabilities, and enhanced environmental interpretation to improve real-world adaptability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。