首例在轨运行零样本视觉语言模型,实现卫星自主智能分析地球影像。
NAVI-Orbital: First In-Orbit Demonstration of a Zero-Shot Vision-Language Model for Autonomous Earth Observation

- 用Gemma 3模型在轨完成图像分类与自然语言描述生成
- 在88.16%准确率下成功识别未见过的地球影像
- 无需微调即可在卫星边缘设备上运行,适合遥感智能处理
随着地球观测数据生成速度超过下行带宽和人工处理能力,星上采集与地面可行动情报之间的差距日益扩大。本文介绍部署于低地球轨道(LEO)航天器上的NAVI-Orbital软件系统。2026年4月16日,NAVI-Orbital实现了迄今作者所知首次在轨演示的视觉语言模型全自主多模态推理。该系统利用本地视觉语言模型(Gemma 3),对每幅捕获场景进行分类,生成内容描述及特征间关系文本,并通过自然语言对话响应操作员提问。系统通过纯英文提示重新任务,而非传统指令序列,由基于图的状态机(LangGraph)协调检测与对话专用代理。地面基准测试(在7,960张精选AID数据集上达88.16%准确率)、模拟平台验证以及实时在轨获取的全新未见地球影像(包括未经校正的YAM-9影像)均表明,在无飞行仪器微调的情况下,通过硬件加速GPU推理,可在卫星级边缘计算机上运行基础模型,实现语义压缩,颠覆传统‘全量采集-全部下传’的带宽模式。
原文摘要 · Abstract (English)
As Earth Observation data generation outpaces downlink bandwidth and human-in-the-loop processing, a widening gap has emerged between onboard collection and actionable ground intelligence. This paper presents NAVI-Orbital, a software system deployed on a Low Earth Orbit (LEO) spacecraft. On April 16, 2026, NAVI-Orbital achieved what is, to the authors' knowledge, the first in-orbit demonstration of a vision-language model performing autonomous multi-modal inference entirely onboard. NAVI-Orbital uses a local vision-language model (Gemma 3) to classify each captured scene, produce a text description of its content and the relationships between its features, and respond to operator follow-up via natural-language dialogue. The system is re-tasked through plain-English prompts in place of conventional command sequences, and is orchestrated by a graph-based state machine (LangGraph) coordinating dedicated agents for detection and dialogue. Results across ground benchmarking (88.16% accuracy on the 7,960-image curated AID benchmark), Flatsat validation, and live in-orbit captures of newly acquired, previously unseen Earth imagery (including uncorrected YAM-9 imagery, processed onboard with hardware-accelerated GPU inference and no fine-tuning for the flight instrument) demonstrate the feasibility of running foundation models on satellite-class edge computers to invert the conventional acquire-then-downlink-everything bandwidth profile through semantic compression of Earth observations in-orbit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。