arXiv:2504.21194cs.CVcs.AI2025-04

构建首个国际空间站航拍图像地理定位基准,提升自动定位准确率至90%。

ISS-Geo142: A Benchmark for Geolocating Astronaut Photography from the International Space Station

  • 采用三类方法:神经网络、SIFT特征匹配与GPT-4视觉推理系统。
  • 新基准涵盖142张图像,最高定位准确率达90%。
  • 适合遥感、计算机视觉与空间信息研究者参考。

本文提出ISS-Geo142,一个用于国际空间站(ISS)宇航员拍摄图像地理定位的精心构建基准。尽管拍摄时的空间站位置已知,但图像所展示的地球具体位置通常未直接标注,导致自动化定位极具挑战。该基准包含142张带元数据和人工标注地理坐标的图像,覆盖多种空间尺度与场景类型。基于此,我们实现并评估三种定位流程:基于VGG16特征与地图兴趣区(AOI)交叉相关性的神经网络方法(NN-Geo),使用滑动窗口特征匹配在拼接高分辨率AOI上的SIFT-Match方法,以及基于具备视觉能力的GPT-4模型的TerraByte系统,可联合推理图像内容与空间站坐标。在本基准上,NN-Geo在评估协议下成功匹配75.52%图像,SIFT-Match在结构丰富场景中表现高精度但计算开销大,而TerraByte建立最强整体基线,约90%图像定位正确,并生成可读地理描述。方法与实验原于2023年开发,本文为修订扩展版,将其置于后续跨视图地理定位及遥感视觉-语言模型进展背景下。总体而言,ISS-Geo142与三类流程共同提供了一个具历史意义、可落地的未来研究基准。

原文摘要 · Abstract (English)

This paper introduces ISS-Geo142, a curated benchmark for geolocating astronaut photography captured from the International Space Station (ISS). Although the ISS position at capture time is known precisely, the specific Earth locations depicted in these images are typically not directly georeferenced, making automated localization non-trivial. ISS-Geo142 consists of 142 images with associated metadata and manually determined geographic locations, spanning a range of spatial scales and scene types. On top of this benchmark, we implement and evaluate three geolocation pipelines: a neural network based approach (NN-Geo) using VGG16 features and cross-correlation over map-derived Areas of Interest (AOIs), a Scale-Invariant Feature Transform based pipeline (SIFT-Match) using sliding-window feature matching on stitched high-resolution AOIs, and TerraByte, an AI system built around a GPT-4 model with vision capabilities that jointly reasons over image content and ISS coordinates. On ISS-Geo142, NN-Geo achieves a match for 75.52\% of the images under our evaluation protocol, SIFT-Match attains high precision on structurally rich scenes at substantial computational cost, and TerraByte establishes the strongest overall baseline, correctly geolocating approximately 90\% of the images while also producing human-readable geographic descriptions. The methods and experiments were originally developed in 2023; this manuscript is a revised and extended version that situates the work relative to subsequent advances in cross-view geo-localization and remote-sensing vision--language models. Taken together, ISS-Geo142 and these three pipelines provide a concrete, historically grounded benchmark for future work on ISS image geolocation.

地理定位空间站影像视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。