构建首个超大规模遥感图像指代分割数据集,推动跨模态理解发展
A Large-Scale Referring Remote Sensing Image Segmentation Dataset and Benchmark
- 提出多尺度指代分割网络,融合细粒度与跨尺度特征
- 在1.5万张高分辨率遥感图上实现领先性能,支持单/多/非目标分割
- 数据集覆盖30+国家、4.9万标注目标,适合遥感与多模态研究者
指代遥感图像分割是一项融合计算机视觉与自然语言处理的复杂挑战。现有数据集在分辨率、场景多样性及类别覆盖方面存在显著局限,制约了分割模型的泛化能力与实际应用。为此,我们推出目前最大最多样化的遥感指代分割数据集NWPU-Refer,包含15,003张高分辨率图像(1024-2048像素),覆盖30多个国家,标注49,745个目标,支持单目标、多目标及非目标分割场景。同时提出专为遥感设计的多尺度指代分割网络MRSNet,引入两个创新模块:(1)同尺度特征交互模块(IFIM),捕捉编码器各阶段的细粒度信息;(2)分层特征交互模块(HFIM),实现跨尺度特征无缝融合,保持空间完整性并增强判别力。在NWPU-Refer上的大量实验表明,MRSNet在多个评估指标上达到当前最优表现,验证了其有效性。数据集与代码已开源:https://github.com/CVer-Yang/NWPU-Refer。
原文摘要 · Abstract (English)
Referring Remote Sensing Image Segmentation is a complex and challenging task that integrates the paradigms of computer vision and natural language processing. Existing datasets for RRSIS suffer from critical limitations in resolution, scene diversity, and category coverage, which hinders the generalization and real-world applicability of refer segmentation models. To facilitate the development of this field, we introduce NWPU-Refer, the largest and most diverse RRSIS dataset to date, comprising 15,003 high-resolution images (1024-2048px) spanning 30+ countries with 49,745 annotated targets supporting single-object, multi-object, and non-object segmentation scenarios. Additionally, we propose the Multi-scale Referring Segmentation Network (MRSNet), a novel framework tailored for the unique demands of RRSIS. MRSNet introduces two key innovations: (1) an Intra-scale Feature Interaction Module (IFIM) that captures fine-grained details within each encoder stage, and (2) a Hierarchical Feature Interaction Module (HFIM) to enable seamless cross-scale feature fusion, preserving spatial integrity while enhancing discriminative power. Extensive experiments conducte on the proposed NWPU-Refer dataset demonstrate that MRSNet achieves state-of-the-art performance across multiple evaluation metrics, validating its effectiveness. The dataset and code are publicly available at https://github.com/CVer-Yang/NWPU-Refer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。