无需训练,用视觉模型补全缺失的高程数据,提升城市数字表面模型精度。
Test-Time Adaptation for Height Completion via Self-Supervised ViT Features and Monocular Foundation Models
- 利用自监督ViT特征与单目深度模型,在测试时通过语义对应传播度量信息。
- 在真实数据上将重建误差降低46%,优于线性拟合和现有单目深度模型。
- 适合需要快速更新、跨域通用的地理信息应用,如城市监测与灾后评估。
精确的数字表面模型(DSMs)对城市监测、环境分析、基础设施管理及变化检测等地理空间应用至关重要。然而,大规模DSM常因采集限制、重建伪影或建成环境变化导致区域缺失或过时。传统高程补全方法依赖空间插值,假设空间连续性,当物体缺失时失效。近年学习方法虽提升重建质量,但通常需针对传感器特定数据集进行监督训练,泛化能力受限。本文提出Prior2DSM,一种完全在测试阶段运行的无训练框架,利用基础模型实现度量高程补全。该方法结合DINOv3的自监督视觉变换器特征与单目深度基础模型,通过语义特征空间对应,将不完整高程先验中的度量信息进行传播。测试时适配(TTA)采用参数高效低秩适配(LoRA)与轻量级多层感知机(MLP),预测空间变化的缩放与偏移参数,将相对深度估计转换为度量高程。实验表明,该方法在多个基准上持续优于基于插值、先验重缩放及当前最优单目深度估计模型。Prior2DSM不仅降低重建误差,保持结构保真度,且实现DSM更新与RGB-DSM联合生成。
原文摘要 · Abstract (English)
Accurate digital surface models (DSMs) are essential for many geospatial applications, including urban monitoring, environmental analyses, infrastructure management, and change detection. However, large-scale DSMs frequently contain incomplete or outdated regions due to acquisition limitations, reconstruction artifacts, or changes in the built environment. Traditional height completion approaches primarily rely on spatial interpolation or which assume spatial continuity and therefore fail when objects are missing. Recent learning-based approaches improve reconstruction quality but typically require supervised training on sensor-specific datasets, limiting their generalization across domains and sensing conditions. We propose Prior2DSM, a training-free framework for metric DSM completion that operates entirely at test time by leveraging foundation models. Unlike previous height completion approaches that require task-specific training, the proposed method combines self-supervised Vision Transformer (ViT) features from DINOv3 with monocular depth foundation models to propagate metric information from incomplete height priors through semantic feature-space correspondence. Test-time adaptation (TTA) is performed using parameter-efficient low-rank adaptation (LoRA) together with a lightweight multilayer perceptron (MLP), which predicts spatially varying scale and shift parameters to convert relative depth estimates into metric heights. Experiments demonstrate consistent improvements over interpolation based methods, prior-based rescaling height approaches, and state-of-the-art monocular depth estimation models. Prior2DSM reduces reconstruction error while preserving structural fidelity, achieving up to a 46% reduction in RMSE compared to linear fitting of MDE, and further enables DSM updating and coupled RGB-DSM generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。