卫星先生成摘要再传数据,省带宽还提速洞察。
Summarize First, Download Later: Onboard VLMs for Bandwidth-Efficient Earth Observation

- 用压缩版视觉语言模型在轨生成自然语言摘要
- 仅在确认关键信息后才下载全分辨率图像,带宽降超80%
- 适合需要快速响应的遥感任务,如火灾监测
现代地球观测卫星搭载越来越先进的传感器,产生海量高分辨率多光谱数据,但下行链路容量仍是关键瓶颈,常导致显著延迟或宝贵观测数据丢失。我们提出“先摘要,后下载”的新范式,利用最新的星载边缘计算和视觉语言模型(VLMs)。卫星不再无差别下传原始图像,而是采用三阶段交互协议:首先上传由量化后的星载VLM生成的简洁自然语言摘要;地面操作员随后发出针对性的视觉问答(VQA)查询,以验证场景相关性(如火灾或海上异常);仅当确认关键信息时,才下载全分辨率图像。该方法将下行链路从被动的数据传输转变为主动的语义感知对话。我们在资源受限的NVIDIA Jetson平台上实现并评估该系统,实验在多种遥感场景中表明,该策略显著降低带宽消耗,同时加快对时效性任务的洞察速度。
原文摘要 · Abstract (English)
Modern Earth observation (EO) satellites carry increasingly advanced sensors that produce vast volumes of high-resolution, multispectral data, yet downlink capacity remains a critical bottleneck -- often causing significant latency or the loss of valuable observations within limited contact windows. We propose a "Summarize First, Download Later" paradigm that exploits recent advances in onboard edge computing and Vision-Language Models (VLMs). Rather than indiscriminately downlinking raw imagery, the system follows a three-phase interaction protocol: the satellite first transmits concise natural language summaries generated by a quantized onboard VLM; ground operators then issue targeted Visual Question Answering (VQA) queries to verify scene relevance (e.g., wildfires or maritime anomalies); and full-resolution images are downloaded only when critical information is confirmed. This transforms the downlink from passive bulk transfer into an active, semantics-aware dialogue. We implement and evaluate the system on a resource-constrained NVIDIA Jetson platform, and experiments on diverse remote sensing scenes show that the proposed strategy substantially reduces bandwidth consumption while accelerating time-to-insight for time-sensitive missions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。