测试大模型在复杂航天任务中的推理能力,发现能想却不会做。
Can LLMs Do Rocket Science? Exploring the Limits of Complex Reasoning with GTOC 12
- 用航天竞赛题评估大模型的多阶段自主规划能力
- 两年间策略可行性得分从9.3升至17.2(满分26)
- 模型懂概念但常因单位错误、边界条件出错而失败
大型语言模型(LLMs)在代码生成和通用推理方面表现优异,但在高维物理约束环境下实现自主多阶段规划的能力仍不明确。本研究通过第12届全球轨道优化竞赛(GTOC 12)这一复杂天体动力学挑战,评估当前AI代理的极限,该任务要求设计大规模小行星采矿方案。我们基于MLE-Bench框架适配轨道力学领域,并采用AIDE架构的智能体自主生成与优化任务方案。为超越单纯有效性判断,引入“大模型作为裁判”方法,依据领域专家制定的评分标准,在五个结构维度上评估战略可行性。对比分析涵盖GPT-4-Turbo至Gemini 2.5 Pro、o3等模型,结果显示:平均战略可行性得分在过去两年间近乎翻倍(由9.3升至17.2/26)。然而,我们识别出策略与执行间的重大能力鸿沟:尽管先进模型具备成熟的概念理解力,能正确设定目标函数与任务架构,却持续在实施中因物理单位不一致、边界条件错误及低效调试循环而失败。结论表明,当前大模型虽具备应对太空科学任务的知识与智能,但仍受限于实现障碍,更像强大的领域协作者而非完全自主工程师。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable proficiency in code generation and general reasoning, yet their capacity for autonomous multi-stage planning in high-dimensional, physically constrained environments remains an open research question. This study investigates the limits of current AI agents by evaluating them against the 12th Global Trajectory Optimization Competition (GTOC 12), a complex astrodynamics challenge requiring the design of a large-scale asteroid mining campaign. We adapt the MLE-Bench framework to the domain of orbital mechanics and deploy an AIDE-based agent architecture to autonomously generate and refine mission solutions. To assess performance beyond binary validity, we employ an "LLM-as-a-Judge" methodology, utilizing a rubric developed by domain experts to evaluate strategic viability across five structural categories. A comparative analysis of models, ranging from GPT-4-Turbo to reasoning-enhanced architectures like Gemini 2.5 Pro, and o3, reveals a significant trend: the average strategic viability score has nearly doubled in the last two years (rising from 9.3 to 17.2 out of 26). However, we identify a critical capability gap between strategy and execution. While advanced models demonstrate sophisticated conceptual understanding, correctly framing objective functions and mission architectures, they consistently fail at implementation due to physical unit inconsistencies, boundary condition errors, and inefficient debugging loops. We conclude that, while current LLMs often demonstrate sufficient knowledge and intelligence to tackle space science tasks, they remain limited by an implementation barrier, functioning as powerful domain facilitators rather than fully autonomous engineers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。