arXiv:2604.04667cs.CVcs.LG2026-04

用束差法让扩散模型实时生成精准航拍深度图。

ZeD-MAP: Bundle Adjustment Guided Zero-Shot Depth Maps for Real-Time Aerial Imaging

  • 用增量式束差法优化航拍帧群,生成一致的三维姿态和稀疏点云
  • 在50米高度下实现水平误差0.87米、垂直误差0.12米的亚米级精度
  • 适合灾备等需实时三维建图的高分辨率无人机应用

从超高清无人机影像中实时重建深度对灾害响应等地理空间任务至关重要,但受限于宽基线视差、大图像尺寸、低纹理或镜面表面、遮挡及严格计算约束。现有零样本扩散模型虽可快速生成稠密深度图且无需特定任务训练,也减少标注数据需求,但其概率推断导致度量准确性与连续帧间不一致。本文提出ZeD-MAP,通过引入基于聚类的增量束差法(BA),将测试时的扩散深度模型转化为类似SLAM的度量一致映射流程。无人机帧被分组为重叠簇,周期性执行BA以获得度量一致的姿态和稀疏3D同名点,并重投影至选定帧作为度量引导,用于扩散深度估计。在约50米高度(地面采样间距约为0.85厘米/像素,每帧覆盖约2,650平方米)使用DLR模块化航空相机系统(MACS)采集的地面标记飞行数据验证表明,本方法实现亚米级精度:水平(XY)方向误差约0.87米,垂直(Z)方向误差0.12米,单帧处理时间保持在1.47至4.91秒之间。结果受人工点云标注噪声影响较小。表明基于束差法的度量引导可达到经典摄影测量一致性,同时显著加速处理,支持实时三维地图生成。

原文摘要 · Abstract (English)

Real-time depth reconstruction from ultra-high-resolution UAV imagery is essential for time-critical geospatial tasks such as disaster response, yet remains challenging due to wide-baseline parallax, large image sizes, low-texture or specular surfaces, occlusions, and strict computational constraints. Recent zero-shot diffusion models offer fast per-image dense predictions without task-specific retraining, and require fewer labelled datasets than transformer-based predictors while avoiding the rigid capture geometry requirement of classical multi-view stereo. However, their probabilistic inference prevents reliable metric accuracy and temporal consistency across sequential frames and overlapping tiles. We present ZeD-MAP, a cluster-level framework that converts a test-time diffusion depth model into a metrically consistent, SLAM-like mapping pipeline by integrating incremental cluster-based bundle adjustment (BA). Streamed UAV frames are grouped into overlapping clusters; periodic BA produces metrically consistent poses and sparse 3D tie-points, which are reprojected into selected frames and used as metric guidance for diffusion-based depth estimation. Validation on ground-marker flights captured at approximately 50 m altitude (GSD is approximately 0.85 cm/px, corresponding to 2,650 square meters ground coverage per frame) with the DLR Modular Aerial Camera System (MACS) shows that our method achieves sub-meter accuracy, with approximately 0.87 m error in the horizontal (XY) plane and 0.12 m in the vertical (Z) direction, while maintaining per-image runtimes between 1.47 and 4.91 seconds. Results are subject to minor noise from manual point-cloud annotation. These findings show that BA-based metric guidance provides consistency comparable to classical photogrammetric methods while significantly accelerating processing, enabling real-time 3D map generation.

无人机深度估计束差法实时建图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。