arXiv:2503.11219cs.CVcs.AI2025-03被引 41

构建百万级无缩放遥感场景数据集,提升精细地理分类精度

MEET: A Million-Scale Dataset for Fine-Grained Geospatial Scene Classification with Zoom-Free Remote Sensing Imagery

  • 提出无缩放遥感图像的场景内布局,中心与辅助场景协同标注
  • 基于MEET数据集,新模型CAT在Swin-Large上提升1.88%平衡准确率
  • 适用于城市功能区划分等实际应用,支持高精度空间上下文建模

利用遥感图像进行精确的细粒度地理场景分类对众多应用至关重要。然而,现有方法通常依赖手动缩放图像以生成典型场景样本,难以满足实际场景中固定分辨率的图像解析需求。为此,我们提出百万级细粒度地理场景分类数据集MEET,包含超过103万张无缩放遥感场景样本,人工标注为80个细粒度类别。在MEET中,每个场景样本采用场景内布局,中心场景作为参考,辅助场景提供关键空间上下文信息。针对场景内分类的新挑战,我们提出上下文感知变压器(CAT),该模型通过学习中心场景与辅助场景间的注意力特征,自适应融合空间上下文,实现精准分类。基于MEET,我们建立了一个全面的基准测试,评估CAT与11个先进基线模型的性能。结果表明,使用Swin-Large骨干网络时,CAT在平衡准确率(BA)上提升1.88%,而使用Swin-Huge骨干网络时更提升7.87%。进一步实验验证了CAT各模块的有效性,并展示了其在城市功能区映射中的实际应用价值。代码与数据集将公开发布于https://jerrywyn.github.io/project/MEET.html。

原文摘要 · Abstract (English)

Accurate fine-grained geospatial scene classification using remote sensing imagery is essential for a wide range of applications. However, existing approaches often rely on manually zooming remote sensing images at different scales to create typical scene samples. This approach fails to adequately support the fixed-resolution image interpretation requirements in real-world scenarios. To address this limitation, we introduce the Million-scale finE-grained geospatial scEne classification dataseT (MEET), which contains over 1.03 million zoom-free remote sensing scene samples, manually annotated into 80 fine-grained categories. In MEET, each scene sample follows a scene-inscene layout, where the central scene serves as the reference, and auxiliary scenes provide crucial spatial context for finegrained classification. Moreover, to tackle the emerging challenge of scene-in-scene classification, we present the Context-Aware Transformer (CAT), a model specifically designed for this task, which adaptively fuses spatial context to accurately classify the scene samples. CAT adaptively fuses spatial context to accurately classify the scene samples by learning attentional features that capture the relationships between the center and auxiliary scenes. Based on MEET, we establish a comprehensive benchmark for fine-grained geospatial scene classification, evaluating CAT against 11 competitive baselines. The results demonstrate that CAT significantly outperforms these baselines, achieving a 1.88% higher balanced accuracy (BA) with the Swin-Large backbone, and a notable 7.87% improvement with the Swin-Huge backbone. Further experiments validate the effectiveness of each module in CAT and show the practical applicability of CAT in the urban functional zone mapping. The source code and dataset will be publicly available at https://jerrywyn.github.io/project/MEET.html.

遥感图像细粒度分类场景上下文数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。