arXiv:2601.11801cs.ROcs.AI2026-01

用视觉语言模型自动设计仿生机器人,省去大量人工干预。

RobotDesignGPT: Automated Robot Design Synthesis using Vision Language Models

  • 输入用户描述和参考图,用大模型自动生成初始设计。
  • 通过视觉反馈机制提升设计质量,减少人工修改次数。
  • 可生成类动物的腿式、飞行动物等兼具美观与运动可行性的结构。

机器人设计是一个复杂的多准则过程,需兼顾用户需求、运动结构和外观造型,通常依赖领域专家经验和大量人力投入。现有方法多为基于规则的,需预先定义语法规则或基础组件模块进行组合设计。本文提出一种新框架 RobotDesignGPT,利用大型预训练视觉语言模型的通用知识与推理能力,实现机器人设计的自动化合成。该框架从简单用户提示和参考图像出发,生成初始设计方案,并引入新颖的视觉反馈机制,显著提升设计质量并减少不必要的手动调整。实验表明,该框架能生成视觉吸引力强且运动学合理的仿生机器人,涵盖腿式动物到飞行生物等多种形态。通过消融实验和用户研究验证了方法的有效性。

原文摘要 · Abstract (English)

Robot design is a nontrivial process that involves careful consideration of multiple criteria, including user specifications, kinematic structures, and visual appearance. Therefore, the design process often relies heavily on domain expertise and significant human effort. The majority of current methods are rule-based, requiring the specification of a grammar or a set of primitive components and modules that can be composed to create a design. We propose a novel automated robot design framework, RobotDesignGPT, that leverages the general knowledge and reasoning capabilities of large pre-trained vision-language models to automate the robot design synthesis process. Our framework synthesizes an initial robot design from a simple user prompt and a reference image. Our novel visual feedback approach allows us to greatly improve the design quality and reduce unnecessary manual feedback. We demonstrate that our framework can design visually appealing and kinematically valid robots inspired by nature, ranging from legged animals to flying creatures. We justify the proposed framework by conducting an ablation study and a user study.

机器人设计视觉语言模型自动合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。