用大模型联合优化机器人形态与奖励,自动设计更优的运动结构。
RoboMoRe: LLM-based Robot Co-design via Joint Optimization of Morphology and Reward
- 通过大模型生成多样且高质量的形态-奖励组合,探索设计空间。
- 在八项任务中超越人工设计和现有方法,无需特定提示或模板。
- 适合对自动化机器人设计感兴趣的研究者与工程师。
机器人协同设计——即联合优化形态与控制策略——仍是机器人领域长期挑战。现有方法常因固定奖励函数而收敛到次优设计,难以适配不同形态的多样化运动模式。本文提出RoboMoRe,一种基于大语言模型(LLM)的框架,将形态与奖励塑造整合进协同优化循环。该框架采用双阶段优化:粗粒度阶段,利用LLM驱动的多样性反射机制生成多样且高质量的形态-奖励对,高效探索其分布;细粒度阶段,通过交替进行LLM引导的奖励与形态梯度更新,迭代优化候选方案。RoboMoRe可实现高效形态及其匹配运动行为的联合优化。实验表明,在不使用任何任务特定提示或预设奖励/形态模板的前提下,其在八项不同任务中显著优于人工设计及现有方法。
原文摘要 · Abstract (English)
Robot co-design, jointly optimizing morphology and control policy, remains a longstanding challenge in the robotics community, where many promising robots have been developed. However, a key limitation lies in its tendency to converge to sub-optimal designs due to the use of fixed reward functions, which fail to explore the diverse motion modes suitable for different morphologies. Here we propose RoboMoRe, a large language model (LLM)-driven framework that integrates morphology and reward shaping for co-optimization within the robot co-design loop. RoboMoRe performs a dual-stage optimization: in the coarse optimization stage, an LLM-based diversity reflection mechanism generates both diverse and high-quality morphology-reward pairs and efficiently explores their distribution. In the fine optimization stage, top candidates are iteratively refined through alternating LLM-guided reward and morphology gradient updates. RoboMoRe can optimize both efficient robot morphologies and their suited motion behaviors through reward shaping. Results demonstrate that without any task-specific prompting or predefined reward/morphology templates, RoboMoRe significantly outperforms human-engineered designs and competing methods across eight different tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。