arXiv:2606.20523cs.CVcs.AI2026-06

构建首个80cm级高分辨率SAR-光学-文本数据集,支持跨模态学习

SARLO-80: Worldwide Slant SAR Language Optic Dataset 80cm

论文配图:SARLO-80: Worldwide Slant SAR Language Optic Dataset 80cm
图 1 · 摘自论文原文
  • 将全球2500景伞形卫星雷达数据重采样至80cm斜距网格,对齐光学影像
  • 生成11.9万组三元组数据,覆盖72国257个地点,含长短三类文本描述
  • 提供完整预处理代码和划分方案,支持原生雷达几何下的多模态研究

多模态基础模型的快速发展得益于大规模光学基准数据集,但合成孔径雷达(SAR)领域仍缺乏可比资源。现有SAR-光学数据集多依赖低分辨率、仅强度的地面范围检测(GRD)产品,未保留复数形式的SAR测量值或原始成像几何,限制了物理驱动的多模态学习。特别是,结合超高清(VHR)SAR SLC、对齐光学影像和自然语言描述的大规模公开数据集依然稀缺。本文基于开源的伞形卫星聚光灯模式数据,以传感器无关复数数据(SICD)形式构建数据集。从约2500个全球场景(VV/HH极化,20厘米至2米原生分辨率)中,通过带限傅里叶变换重采样,将所有雷达数据标准化至80厘米斜距网格,并将图像切分为1024×1024像素块。每个SAR块对应一个高分辨率光学块,利用局部坐标对应关系将其映射至SAR网格实现像素级对齐。每条样本生成三种长度的文本描述(短/中/长),以支持视觉-语言训练与评估。数据集共包含119,566组三元组(复数及幅度斜距SAR块、对齐光学块、自然语言描述),覆盖72个国家的257个地点,涵盖多种地表类型与基础设施。我们提供固定的训练/验证/测试划分及完整预处理与基线代码,支持在原生雷达几何下进行跨模态检索与条件生成的可复现基准。数据集已公开发布于Hugging Face Hub:https://huggingface.co/datasets/ONERA/SARLO-80。

原文摘要 · Abstract (English)

Multimodal foundation models have advanced rapidly thanks to large optical benchmarks, but comparable resources for synthetic aperture radar (SAR) remain limited. Existing SAR--optical datasets largely rely on low-resolution, intensity-only Ground Range Detected~(GRD) products and do not preserve complex-valued SAR measurements or native acquisition geometry, which restricts physically grounded multimodal learning. In particular, large-scale public datasets combining very-high-resolution (VHR) SAR SLC, aligned optical imagery, and natural-language descriptions are still lacking. We present a VHR SAR--optical--text dataset built from open-access Umbra spotlight acquisitions distributed as Sensor Independent Complex Data (SICD). From around 2,500 worldwide scenes (VV/HH, 20cm--2m native resolution), we standardize all SAR data to an 80cm slant-range grid via band-limited FFT resampling and tile the imagery into 1024 by 1024 patches. For each SAR patch, we retrieve a high-resolution optical tile and warp it into the SAR grid using local coordinate correspondences for local pixel-level alignment. We further generate three caption variants (SHORT/MID/LONG) per sample to support vision--language training and evaluation. Our dataset contains 119,566 triplets (complex and amplitude slant-range SAR patch, aligned optical patch, natural-language description) covering 257 locations across 72 countries and a broad range of land types and infrastructures. We release fixed train/validation/test splits and the full preprocessing and baseline code to enable reproducible benchmarks for multimodal alignment on cross-modal retrieval and conditional generation in native SAR geometry. The dataset is publicly available on the Hugging Face Hub at https://huggingface.co/datasets/ONERA/SARLO-80.

SAR数据集多模态遥感文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。