用强化学习训练的模型能设计出比人类还好的火箭。
LLMs for Engineering: Teaching Models to Design High Powered Rockets
- 用强化学习让小模型迭代优化火箭设计
- 70亿参数模型在登月任务中超越人类专家
- 为物理工程提供可自动优化的智能工具
大型语言模型(LLMs)已深刻改变软件工程,但在物理工程领域的应用仍不充分。本文通过RocketBench基准,将LLMs与高保真火箭仿真系统连接,评估其在高功率火箭设计中的表现。测试任务包括目标高度优化和精准着陆挑战,逐步提升复杂度。结果显示,当前顶尖LLMs虽具备基础工程知识,但无法有效根据仿真反馈迭代改进,最终性能低于人类水平。然而,在引入强化学习(RL)后,一个70亿参数模型不仅超越了现有主流基础模型,更在精度着陆任务中超过人类专家。该研究证明,经过强化学习训练的LLM可成为复杂工程优化的有效工具,有望推动工程领域从软件之外的广泛变革。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have transformed software engineering, but their application to physical engineering domains remains underexplored. This paper evaluates LLMs' capabilities in high-powered rocketry design through RocketBench, a benchmark connecting LLMs to high-fidelity rocket simulations. We test models on two increasingly complex design tasks: target altitude optimization and precision landing challenges. Our findings reveal that while state-of-the-art LLMs demonstrate strong baseline engineering knowledge, they struggle to iterate on their designs when given simulation results and ultimately plateau below human performance levels. However, when enhanced with reinforcement learning (RL), we show that a 7B parameter model outperforms both SoTA foundation models and human experts. This research demonstrates that RL-trained LLMs can serve as effective tools for complex engineering optimization, potentially transforming engineering domains beyond software development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。