用状态空间模型统一解决卫星图建筑分割与高度估计难题
BuildMamba: A Visual State-Space Based Model for Multi-Task Building Segmentation and Height Estimation from Satellite Images
- 基于视觉状态空间模型实现线性时间全局建模
- 在DFC23上达0.93的分割精度和1.77米高度误差
- 适合城市三维重建、遥感分析等实际应用
从单视角RGB卫星图像中准确进行建筑分割与高度估计是城市分析的基础,但受结构多样性及全局上下文建模高计算成本影响,仍属未解难题。现有方法多沿用单目深度网络,常出现边界模糊和高层建筑系统性低估问题。为此,我们提出BuildMamba,一种统一多任务框架,利用视觉状态空间模型的线性时间全局建模能力。为增强结构耦合性与计算效率,设计三个模块:动态空间重校准的Mamba注意力模块,通过门控状态空间扫描实现多尺度特征聚合的Spatial-Aware Mamba-FPN,以及结合语义先验抑制高度伪影的Mask-Aware Height Refinement模块。大量实验表明,BuildMamba在三个基准上建立新性能上限。尤其在DFC23基准上,分割交并比达0.93,高度估计均方根误差低至1.77米,较最先进方法提升0.82米。仿真结果验证了模型在大规模3D城市重建中的卓越鲁棒性与可扩展性。
原文摘要 · Abstract (English)
Accurate building segmentation and height estimation from single-view RGB satellite imagery are fundamental for urban analytics, yet remain ill-posed due to structural variability and the high computational cost of global context modeling. While current approaches typically adapt monocular depth architectures, they often suffer from boundary bleeding and systematic underestimation of high-rise structures. To address these limitations, we propose BuildMamba, a unified multi-task framework designed to exploit the linear-time global modeling of visual state-space models. Motivated by the need for stronger structural coupling and computational efficiency, we introduce three modules: a Mamba Attention Module for dynamic spatial recalibration, a Spatial-Aware Mamba-FPN for multi-scale feature aggregation via gated state-space scans, and a Mask-Aware Height Refinement module using semantic priors to suppress height artifacts. Extensive experiments demonstrate that BuildMamba establishes a new performance upper bound across three benchmarks. Specifically, it achieves an IoU of 0.93 and RMSE of 1.77~m on DFC23 benchmark, surpassing state-of-the-art by 0.82~m in height estimation. Simulation results confirm the model's superior robustness and scalability for large-scale 3D urban reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。