构建首个面向遥感的通用视觉模型,支持多任务感知。
RemoteSAM: Towards Segment Anything for Earth Observation
- 用自动数据引擎生成27万组图文掩码,覆盖超多语义类别。
- 单模型统一处理分类、检测、分割等任务,性能超越现有模型。
- 适合遥感分析、地理信息研究者使用,开源可复现。
我们旨在开发一个鲁棒且灵活的地球观测视觉基础模型,具备识别与定位多样视觉目标的能力,并兼容不同任务场景所需的多种输入输出接口。现有系统难以满足这些需求,因它们通常基于窄域数据训练,采用任务专用架构,语义覆盖有限。本研究从数据与建模两方面突破:首先提出自动数据引擎,相比人工标注或规则方法更具可扩展性,成功构建迄今最大规模同类数据集,包含27万张图像-文本-掩码三元组,涵盖前所未有的多样化语义类别与属性描述。在此数据基础上,我们提出以指代表达分割为核心的任务统一范式,仅用单一模型即可高效处理分类、检测、分割、定位等多种视觉感知任务,无需任务特定头。结合数据与建模创新,我们推出RemoteSAM,其在多个地球观测感知基准上达到新SOTA,显著优于Falcon、GeoChat和LHRS-Bot等现有基础模型,同时效率更高。模型与数据已公开于https://github.com/1e12Leon/RemoteSAM。
原文摘要 · Abstract (English)
We aim to develop a robust yet flexible visual foundation model for Earth observation. It should possess strong capabilities in recognizing and localizing diverse visual targets while providing compatibility with various input-output interfaces required across different task scenarios. Current systems cannot meet these requirements, as they typically utilize task-specific architecture trained on narrow data domains with limited semantic coverage. Our study addresses these limitations from two aspects: data and modeling. We first introduce an automatic data engine that enjoys significantly better scalability compared to previous human annotation or rule-based approaches. It has enabled us to create the largest dataset of its kind to date, comprising 270K image-text-mask triplets covering an unprecedented range of diverse semantic categories and attribute specifications. Based on this data foundation, we further propose a task unification paradigm that centers around referring expression segmentation. It effectively handles a wide range of vision-centric perception tasks, including classification, detection, segmentation, grounding, etc, using a single model without any task-specific heads. Combining these innovations on data and modeling, we present RemoteSAM, a foundation model that establishes new SoTA on several earth observation perception benchmarks, outperforming other foundation models such as Falcon, GeoChat, and LHRS-Bot with significantly higher efficiency. Models and data are publicly available at https://github.com/1e12Leon/RemoteSAM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。