用扩散模型联合优化相机位姿与平面布局,提升宽基线全景图重建精度。
BADGR: Bundle Adjustment Diffusion Conditioned by GRadients for Wide-Baseline Floor Plan Reconstruction
- 基于单步莱文伯格-马夸尔特优化器的密集输出,引导扩散模型生成结构化布局
- 在不同输入密度下显著优于当前最优方法,重投影误差大幅降低
- 仅需2D平面数据训练,支持多样输入密度,适合实际场景应用
从宽基线RGB全景图中精确重建相机位姿与平面布局是尚未解决的难题。本文提出BADGR,一种新型扩散模型,通过联合重构与捆绑调整(BA)从粗略状态精炼位姿与布局,利用数十张不同密度图像的1D墙面预测结果。不同于传统引导式扩散模型,BADGR以单步莱文伯格-马夸尔特(LM)优化器的密集实体输出为条件,训练时同时预测相机与墙体位置,并最小化重投影误差以保证视图一致性。去噪扩散过程生成布局的目标,补充了BA优化,引入额外的已学习布局结构约束(如墙邻接、共线性),结合全局上下文缓解密集边界观测带来的误差。该方法仅在2D平面数据上训练,简化数据获取,支持鲁棒增强和多种输入密度。实验验证表明,其在不同输入密度下均显著超越当前最优方法。
原文摘要 · Abstract (English)
Reconstructing precise camera poses and floor plan layouts from wide-baseline RGB panoramas is a difficult and unsolved problem. We introduce BADGR, a novel diffusion model that jointly performs reconstruction and bundle adjustment (BA) to refine poses and layouts from a coarse state, using 1D floor boundary predictions from dozens of images of varying input densities. Unlike a guided diffusion model, BADGR is conditioned on dense per-entity outputs from a single-step Levenberg Marquardt (LM) optimizer and is trained to predict camera and wall positions while minimizing reprojection errors for view-consistency. The objective of layout generation from denoising diffusion process complements BA optimization by providing additional learned layout-structural constraints on top of the co-visible features across images. These constraints help BADGR to make plausible guesses on spatial relations which help constrain pose graph, such as wall adjacency, collinearity, and learn to mitigate errors from dense boundary observations with global contexts. BADGR trains exclusively on 2D floor plans, simplifying data acquisition, enabling robust augmentation, and supporting variety of input densities. Our experiments and analysis validate our method, which significantly outperforms the state-of-the-art pose and floor plan layout reconstruction with different input densities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。