用深度先验融合影像与高程数据,实现城市2D语义和3D高度变化的联合检测。
DPG-CD: Depth-Prior-Guided Cross-Modal Joint 2D-3D Change Detection

- 引入估计深度先验缓解影像与高程图的模态差异
- 多阶段跨时序跨模态融合提升变化感知能力
- 支持高频监测与应急响应,适合城市演变分析
城市空间演化不仅体现为水平扩张,还表现为垂直结构变化。因此,联合捕捉2D语义变化与3D高程变化对城市形态分析和应急管理至关重要。实际应用中,3D数据获取成本高且难频繁更新。采用灾前数字表面模型(DSM)与灾后影像的多时相跨模态输入,是实现高频城市监测、灾害评估与应急响应中3D变化检测的可行方案。然而,该设定仍具挑战性,因影像与DSM存在显著的光谱-几何表征差异,模态差异可能被误认为真实变化。可靠的变化检测需有效融合多时相数据的语义与几何特征。本文提出DPG-CD框架,通过在影像中引入估计深度先验以缩小与DSM的模态差距,并设计门控融合机制,选择性注入几何线索同时保留判别性光谱特征。随后采用多阶段跨时序跨模态特征融合架构提取变化感知特征。最终通过多任务解码器联合预测2D语义变化与3D高程变化,并辅以辅助的DSM重建任务以增强结构一致性与高程估计精度。在两个公开数据集Hi-BCD、3DCD及新构建的NYC-MMCD上的实验表明,DPG-CD在2D与3D变化检测任务上均优于现有最先进方法。
原文摘要 · Abstract (English)
Urban spatial evolution is manifested not only through horizontal expansion but also through vertical structural changes. Consequently, jointly capturing 2D semantic changes and 3D height changes is essential for urban morphology analysis and emergency management. In practical scenarios, collecting 3D observations is often constrained by high acquisition costs and the inability to support frequent updates. The multi-temporal cross-modal input consisting of pre-event Digital Surface Model (DSM) and post-event imagery provides a practical solution for 3D change detection in high-frequency urban monitoring, disaster assessment, and emergency response scenarios. However, this setting remains challenging as imagery and DSM data exhibit significant spectral-geometric representation gaps. Moreover, modality differences may be confused with actual changes, and robust change detection requires effective fusion of semantic and geometric features from multi-temporal data. In this paper, we propose DPG-CD, a depth-prior-guided multi-temporal cross-modal fusion framework for joint 2D semantic and 3D height change detection. Specifically, an estimated depth prior is introduced into the imagery to mitigate the modality gap with DSM. A gated fusion mechanism then selectively injects geometric cues from depth prior while preserving discriminative spectral representations. Subsequently, a multi-stage cross-temporal cross-modal feature fusion architecture is employed to extract change-aware features. Finally, a multi-task decoder jointly predicts 2D semantic changes and 3D height changes, complemented by an auxiliary DSM prediction task to improve structural consistency and height estimation accuracy. Experiments on two public datasets, Hi-BCD and 3DCD, and a new dataset, NYC-MMCD, demonstrate that DPG-CD outperforms state-of-the-art methods on both 2D and 3D change detection tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。