通过形状关联联合估计建筑高度与轮廓,提升城市建模精度。
Morphology-Guided Cross-Task Coupling for Joint Building Height and Footprint Estimation

- 用轮廓信息引导高度预测,实现跨任务特征交互
- 高度误差降低0.24米,轮廓精度保持0.80的高水平
- 方法不依赖输入分辨率,适合多源遥感数据
建筑高度(BH)和建筑轮廓(BF)共同描述城市空间结构,是城市气候、灾害风险与人口分布模型的关键输入。二者通过容积率(FAR)约束耦合,但现有遥感方法通常独立处理。本文提出MorphoFormer框架,包含两个互补机制:(i) BF引导的任务解码器(BGTD),通过轮廓衍生的形态上下文对高度分支进行交叉注意力门控;(ii) 形态一致性损失(MCL),以轮廓推导的高度代理监督真实高度,间接强制轮廓特征编码高度相关结构。模型采用单阶段Swin主干,融合哨兵1号雷达、哨兵2号多光谱和数字高程模型数据,在54个城市的地理区块划分数据集上训练评估。相比同感受野的Swin-MTL基线,MorphoFormer将高度测试均方根误差从3.39降至3.15米(R²由0.62升至0.67),轮廓R²稳定在0.80。消融实验显示,移除BGTD使误差上升0.11米,移除MCL也上升0.11米,剩余约0.02米处于编码器波动噪声范围内。两机制作用于跨任务表示而非像素,无固有分辨率依赖。
原文摘要 · Abstract (English)
Building height (BH) and building footprint (BF) jointly describe the vertical and horizontal extent of the built environment and are required inputs for urban climate, disaster-risk, and population-mapping models. The two parameters are coupled through floor-area-ratio (FAR) constraints, yet remote-sensing approaches typically treat them as independent regression targets. We argue that explicitly encoding this cross-task coupling is more impactful than further refining individual encoders, and propose MorphoFormer, a joint BH/BF estimation framework built around two complementary mechanisms: (i) a BF-Guided Task Decoder (BGTD) that gates the height branch via cross-attention on a footprint-derived morphology context, and (ii) a Morphology Consistency Loss (MCL) that supervises a height-from-footprint surrogate against the ground-truth BH, indirectly forcing the BF feature to encode height-correlated structure. The encoder is a single-stage Swin backbone fed by Sentinel-1 SAR, Sentinel-2 multispectral, and DEM inputs, trained and evaluated on a geo-blocked split of 54 cities. Against a Swin-MTL baseline at identical receptive field, MorphoFormer reduces BH test RMSE from 3.39 to 3.15 m (R^2 improves 0.62 -> 0.67) with BF R^2 stable at 0.80. Controlled ablations at identical capacity attribute most of this 0.24 m improvement to the two proposed mechanisms: removing BGTD raises BH RMSE by 0.11 m and removing MCL raises it by 0.11 m, with the residual approximately 0.02 m falling within the noise floor of encoder-side variations. Because both mechanisms act on cross-task representations rather than pixels, the design carries no intrinsic dependence on input resolution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。