arXiv:2603.11804cs.CVcs.LG2026-03被引 2

用开放地图数据自动生成遥感图文标注,无需人工或大模型教师。

OSMDA: OpenStreetMap-based Domain Adaptation for Remote Sensing VLMs

  • 用航拍图配渲染的开放地图,让模型自动生成带地理元数据的描述。
  • 在六项零样本和五项分布内任务上显著提升性能,超越九个基线方法。
  • 训练成本远低于依赖大模型教师的方法,适合资源有限的研究者。

遥感视觉语言模型(VLMs)依赖领域特定的图像-文本标注,但卫星与航拍图像的高质量标注稀缺且昂贵。现有伪标签方法通过蒸馏大模型知识缓解此问题,但依赖大模型教师导致成本高、可扩展性差,且性能受限于教师上限。本文提出OSMDA:一种自包含的领域适应框架,摆脱对外部教师的依赖。核心思路是利用具备能力的基础VLM作为自身标注引擎:将航拍图像与渲染的开放街景地图(OSM)瓦片配对,借助模型的光学字符识别和图表理解能力,生成富含OSM丰富辅助元数据的描述。随后仅使用卫星图像对该语料库进行微调,得到无需人工标注、也无需外部强教师的OSMDA-VLM。我们在六个零样本和五个分布内基准上进行了全面评估,结果表明OSMDA带来显著性能提升。与九种竞争基线对比显示,本方法整体表现更优,且训练成本大幅降低。这些结果表明,基于强基础模型与众包地理数据对齐,是实现遥感领域适配的一条可行且可扩展路径。数据集与模型权重将在论文接受后公开。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) adapted to remote sensing rely heavily on domain-specific image-text supervision, yet high-quality annotations for satellite and aerial imagery remain scarce and expensive to produce. Prevailing pseudo-labeling pipelines address this gap by distilling knowledge from large frontier models, but this dependence on large teachers is costly, limits scalability, and caps achievable performance at the ceiling of the teacher. We propose OSMDA: a self-contained domain adaptation framework that eliminates this dependency. Our key insight is that a capable base VLM can serve as its own annotation engine: by pairing aerial images with rendered OpenStreetMap (OSM) tiles, we leverage optical character recognition and chart comprehension capabilities of the model to generate captions enriched by OSM's vast auxiliary metadata. The model is then fine-tuned on the resulting corpus with satellite imagery alone, yielding OSMDA-VLM, a domain-adapted VLM that requires no manual labeling and no stronger external VLM teacher. We conduct exhaustive evaluations spanning six zero-shot and five in-distribution benchmarks across vision-language tasks, where OSMDA leads to substantial improvement. We further compare against nine competitive baselines, demonstrating that our method achieves superior overall performance, while being substantially cheaper to train than teacher-dependent alternatives. These results suggest that, given a strong foundation model, alignment with crowd-sourced geographic data is a practical and scalable path towards remote sensing domain adaptation. Dataset and model weights will be made publicly available upon acceptance.

遥感域适应自监督开放地图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。