用视觉语言模型当太空任务操作员,能看图做决策
Visual Language Models as Operator Agents in the Space Domain
- 让VLM通过截图理解界面,自动完成复杂轨道操作
- 在模拟中表现媲美传统方法,真实硬件也能诊断卫星
- 适合航天自动化、人机协同研究者参考
本文探索视觉语言模型(VLMs)作为太空领域操作代理的应用,涵盖软件与硬件两种操作范式。基于大语言模型(LLMs)及其多模态扩展的进展,研究VLM如何提升太空任务中的自主控制与决策能力。在软件层面,将VLM引入Kerbal Space Program Differential Games(KSPDG)仿真环境,使代理可通过解读图形用户界面截图执行复杂轨道机动。在硬件层面,将VLM与带摄像头的机器人系统结合,用于检查和诊断真实空间物体(如卫星)。结果表明,VLM能有效处理视觉与文本数据,生成上下文恰当的操作指令,在仿真任务中表现可比传统方法及非多模态LLM,并展现出在实际应用中的潜力。
原文摘要 · Abstract (English)
This paper explores the application of Vision-Language Models (VLMs) as operator agents in the space domain, focusing on both software and hardware operational paradigms. Building on advances in Large Language Models (LLMs) and their multimodal extensions, we investigate how VLMs can enhance autonomous control and decision-making in space missions. In the software context, we employ VLMs within the Kerbal Space Program Differential Games (KSPDG) simulation environment, enabling the agent to interpret visual screenshots of the graphical user interface to perform complex orbital maneuvers. In the hardware context, we integrate VLMs with robotic systems equipped with cameras to inspect and diagnose physical space objects, such as satellites. Our results demonstrate that VLMs can effectively process visual and textual data to generate contextually appropriate actions, competing with traditional methods and non-multimodal LLMs in simulation tasks, and showing promise in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。