构建首个遥感3D感知综合基准,推动地理AI发展
RS3DBench: A Comprehensive Benchmark for 3D Spatial Perception in Remote Sensing
- 构建5.5万对遥感图像与像素级深度图的对齐数据集
- 提出基于稳定扩散的遥感深度估计模型,性能达当前最优
- 适用于遥感3D视觉、地理AI研究者及多模态模型开发者
本文提出一个新型基准RS3DBench,旨在推动通用大规模遥感3D视觉模型的发展。现有遥感数据集普遍存在深度信息不完整或深度与图像对齐不精确的问题。为此,我们构建了包含54,951对遥感图像与像素级对齐深度图的数据集,并附有对应文本描述,覆盖广泛地理场景。该数据集可用于训练与评估遥感图像空间理解任务中的3D视觉感知模型。此外,我们基于稳定扩散模型开发了一种遥感深度估计方法,利用其多模态融合能力,在本数据集上实现当前最优性能。本工作致力于推动3D视觉感知模型与地理人工智能在遥感领域的进步。数据集、模型与代码将公开于https://rs3dbench.github.io。
原文摘要 · Abstract (English)
In this paper, we introduce a novel benchmark designed to propel the advancement of general-purpose, large-scale 3D vision models for remote sensing imagery. While several datasets have been proposed within the realm of remote sensing, many existing collections either lack comprehensive depth information or fail to establish precise alignment between depth data and remote sensing images. To address this deficiency, we present a visual Benchmark for 3D understanding of Remotely Sensed images, dubbed RS3DBench. This dataset encompasses 54,951 pairs of remote sensing images and pixel-level aligned depth maps, accompanied by corresponding textual descriptions, spanning a broad array of geographical contexts. It serves as a tool for training and assessing 3D visual perception models within remote sensing image spatial understanding tasks. Furthermore, we introduce a remotely sensed depth estimation model derived from stable diffusion, harnessing its multimodal fusion capabilities, thereby delivering state-of-the-art performance on our dataset. Our endeavor seeks to make a profound contribution to the evolution of 3D visual perception models and the advancement of geographic artificial intelligence within the remote sensing domain. The dataset, models and code will be accessed on the https://rs3dbench.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。