arXiv:2606.15427cs.LGcs.AI2026-06被引 1

用自然语言提示让卫星在轨扩展识别能力,无需更新模型权重。

Post-Launch Capability Expansion of Vision-Language Models via Prompting for On-Orbit Spacecraft Inspection

论文配图:Post-Launch Capability Expansion of Vision-Language Models via Prompting for On-Orbit Spacecraft Inspection
图 1 · 摘自论文原文
  • 通过自然语言提示激活视觉-语言模型,实现星载系统上线后语义扩展。
  • 大部件识别准确率超60%,小部件如天线、推进器仍难定位(<25%)。
  • 结构化提示可提升82%性能,适合在轨快速部署新检测任务。

星载感知系统通常在发射前部署模型,此后更新权重或扩展标签集在操作上极不现实。尽管监督模型可在发射前集成,但在轨新增语义能力需重新训练并上传参数。本文研究是否可通过提示驱动的视觉-语言模型实现发射后语义扩展,使新卫星组件可通过自然语言提示指定,而无需修改星上权重。我们在129张未见过的卫星图像测试集上,采用严格冻结、单次推理协议评估零样本实例分割性能。在固定全局阈值且无后处理条件下,SAM3在[email protected][email protected]:0.95分别达到0.385和0.267。性能显著依赖规模:大型结构如卫星本体([email protected]=0.639)和太阳翼([email protected]=0.598)定位可靠,而小型附件如天线([email protected]=0.221)和推进器([email protected]=0.081)仍难以识别。提示形式影响性能,结构化提示(含空间与几何描述)相比简短类别名提示最高可提升82%。模型运行在当前嵌入式GPU的内存与计算容量内,表明提示驱动的语义扩展对主要卫星结构具有实际可行性,但对细粒度组件在轨零样本定位仍存在挑战。

原文摘要 · Abstract (English)

Spaceborne inspection systems often deploy perception models prior to launch, after which updating model weights or expanding fixed label sets becomes operationally impractical. While supervised models can be integrated pre-flight, adding new semantic capabilities in orbit requires retraining and re-uploading parameters. We investigate whether prompt-driven vision--language models can enable post-launch semantic expansion, allowing new spacecraft components to be specified via natural-language prompts without modifying onboard weights. We evaluate zero-shot instance segmentation of spacecraft components under a strictly frozen, single-pass inference protocol on a test set of $129$ images of previously unseen satellites. Under fixed global thresholds and no post-processing, SAM3 achieves $0.385$ mAP@$0.5$ and $0.267$ mAP@$0.5{:}0.95$. Performance is strongly scale-dependent: large structural elements like spacecraft bodies ($0.639$ AP@$0.50$) and solar arrays ($0.598$ AP@$0.5$) localize reliably, while relatively small appendages like antennas ($0.221$ AP@$0.5$) and thrusters ($0.081$ AP@$0.5$) remain difficult. Prompt formulation influences performance, with structured prompts incorporating spatial and geometric descriptors yielding up to $82%$ improvement over short category-name prompts. The model operates within the memory and compute envelope of contemporary embedded GPUs, suggesting prompt-driven grounding can provide a practical mechanism for post-launch semantic extension of dominant spacecraft structures while highlighting limitations of zero-shot localization for fine-scale components under orbital domain shift.

视觉语言在轨扩展提示工程卫星检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。