用多智能体辩论优化机器人形态与奖励,提升设计效率。
Debate2Create: Robot Co-design via Multi-Agent LLM Debate
- 设计与控制智能体通过辩论循环迭代优化,结合物理仿真评估。
- 在5个基准任务中表现最优,最高提升达9倍,迭代辩论增效18%-35%。
- 适合对机器人联合设计感兴趣的算法与工程研究者。
我们提出 Debate2Create(D2C),一个基于多智能体大模型的框架,将机器人协同设计建模为基于物理评估的结构化、迭代式辩论。设计智能体与控制智能体通过论点-反论点-综合循环协作,特定准则的LLM裁判提供多目标反馈以引导探索。在五个MuJoCo运动基准上,D2C在所有对比的基于LLM及黑盒基线中取得最高默认归一化得分,其中Ant任务提升达3.2倍,Swimmer任务接近9倍。迭代辩论相比计算量相当的零样本生成,性能提升18%-35%,且生成的奖励函数在4/5任务中可迁移至默认形态。结果表明,在固定拓扑、单候选强化学习协议下,结构化的、仿真驱动的多智能体交互是联合形态-奖励优化的有效机制。
原文摘要 · Abstract (English)
We introduce Debate2Create (D2C), a multi-agent LLM framework that formulates robot co-design as structured, iterative debate grounded in physics-based evaluation. A design agent and control agent engage in a thesis-antithesis-synthesis loop, while criterion-specific LLM judges provide multi-objective feedback to steer exploration. Across five MuJoCo locomotion benchmarks, D2C achieves the highest default-normalized score among the evaluated LLM-based and black-box baselines, with gains up to 3.2x on Ant and nearly 9x on Swimmer. Iterative debate yields 18-35% gains over compute-matched zero-shot generation, and D2C-generated rewards transfer to default morphologies in 4/5 tasks. These results suggest that structured, simulator-grounded multi-agent interaction is a useful mechanism for joint morphology-reward optimization under a fixed-topology, per-candidate-RL protocol. Project page: debate2create.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。