构建首个面向真实场景的遥感图像分割基准,提升模型泛化能力。
Towards Realistic Open-Vocabulary Remote Sensing Segmentation: Benchmark and Baseline
- 提出新基准OVRSISBenchV2,覆盖128类、17万张遥感图,涵盖建筑、道路等任务
- 设计正向激励噪声机制,使模型在训练中拓展视觉-文本特征空间,增强迁移性
- 适合遥感应用研究者、需要开放词汇分割的复杂场景开发者
开放词汇遥感图像分割(OVRSIS)因数据集分散、训练多样性不足及缺乏反映真实地理应用需求的评估基准而发展受限。此前我们提出的OVRSISBenchV1建立了初步跨数据集评估协议,但范围有限,难以衡量真实世界中的开放域泛化能力。为此,本文提出大规模、面向应用的全新基准OVRSISBenchV2。我们首先构建了约9.5万张图像-掩码对的平衡数据集OVRSIS95K,覆盖35种常见语义类别和多样化遥感场景;在此基础上,结合10个下游数据集,构建出包含17万张图像、128个语义类别的OVRSISBenchV2,显著提升了场景多样性、语义覆盖广度与评估难度。除标准开放词汇分割外,还引入建筑物提取、道路提取和洪水检测等下游任务协议,更贴近真实地理信息应用需求与复杂部署环境。同时,我们提出基线方法Pi-Seg,通过可学习且语义引导的正向激励噪声机制,在训练中扩大视觉-文本特征空间,提升模型迁移能力。在OVRSISBenchV1、OVRSISBenchV2及下游任务上的大量实验表明,Pi-Seg表现强劲且稳定,尤其在更具挑战性的OVRSISBenchV2上优势明显。结果凸显了真实基准设计的重要性及基于扰动的迁移策略的有效性。代码与数据集已开源。
原文摘要 · Abstract (English)
Open-vocabulary remote sensing image segmentation (OVRSIS) remains underexplored due to fragmented datasets, limited training diversity, and the lack of evaluation benchmarks that reflect realistic geospatial application demands. Our previous \textit{OVRSISBenchV1} established an initial cross-dataset evaluation protocol, but its limited scope is insufficient for assessing realistic open-world generalization. To address this issue, we propose \textit{OVRSISBenchV2}, a large-scale and application-oriented benchmark for OVRSIS. We first construct \textbf{OVRSIS95K}, a balanced dataset of about 95K image--mask pairs covering 35 common semantic categories across diverse remote sensing scenes. Built upon OVRSIS95K and 10 downstream datasets, OVRSISBenchV2 contains 170K images and 128 categories, substantially expanding scene diversity, semantic coverage, and evaluation difficulty. Beyond standard open-vocabulary segmentation, it further includes downstream protocols for building extraction, road extraction, and flood detection, thereby better reflecting realistic geospatial application demands and complex deployment scenarios. We also propose \textbf{Pi-Seg}, a baseline for OVRSIS. Pi-Seg improves transferability through a \textbf{positive-incentive noise} mechanism, where learnable and semantically guided perturbations broaden the visual-text feature space during training. Extensive experiments on OVRSISBenchV1, OVRSISBenchV2, and downstream tasks show that Pi-Seg delivers strong and consistent results, particularly on the more challenging OVRSISBenchV2 benchmark. Our results highlight both the importance of realistic benchmark design and the effectiveness of perturbation-based transfer for OVRSIS. The code and datasets are available at \href{https://github.com/LiBingyu01/Pi-Seg}{LiBingyu01/Pi-Seg}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。