arXiv:2502.16779cs.CVcs.AI2025-02ICLR被引 5

用预训练模型直接从多视角图像重建房间布局,省去繁琐步骤

Unposed Sparse Views Room Layout Reconstruction in the Age of Pretrain Model

  • 基于DUSt3R框架,微调后直接估计房间结构平面
  • 在合成数据上超越现有方法,在真实场景中也表现稳健
  • 适合需要快速准确房间布局的三维重建应用

从多视角图像进行房间布局估计因多视图几何复杂性而研究不足,传统方法需相机内参外参估计、图像匹配和三角化等多步流程。近期3D基础模型如DUSt3R已将重建范式转向端到端单步方法。为此,本文提出Plane-DUSt3R,一种基于DUSt3R的多视角房间布局估计新方法。该方法在Structure3D数据集上微调,采用改进目标函数以估计结构平面。通过生成均匀简洁的结果,仅需一次后处理与2D检测即可完成布局估计。相比依赖单视角或全景图的方法,Plane-DUSt3R扩展至多视角设置,提供简化、端到端的解决方案,减少误差累积。实验表明,其在合成数据上超越当前最优方法,并在不同风格的真实图像(如卡通)中仍具鲁棒性与有效性。

原文摘要 · Abstract (English)

Room layout estimation from multiple-perspective images is poorly investigated due to the complexities that emerge from multi-view geometry, which requires muti-step solutions such as camera intrinsic and extrinsic estimation, image matching, and triangulation. However, in 3D reconstruction, the advancement of recent 3D foundation models such as DUSt3R has shifted the paradigm from the traditional multi-step structure-from-motion process to an end-to-end single-step approach. To this end, we introduce Plane-DUSt3R, a novel method for multi-view room layout estimation leveraging the 3D foundation model DUSt3R. Plane-DUSt3R incorporates the DUSt3R framework and fine-tunes on a room layout dataset (Structure3D) with a modified objective to estimate structural planes. By generating uniform and parsimonious results, Plane-DUSt3R enables room layout estimation with only a single post-processing step and 2D detection results. Unlike previous methods that rely on single-perspective or panorama image, Plane-DUSt3R extends the setting to handle multiple-perspective images. Moreover, it offers a streamlined, end-to-end solution that simplifies the process and reduces error accumulation. Experimental results demonstrate that Plane-DUSt3R not only outperforms state-of-the-art methods on the synthetic dataset but also proves robust and effective on in the wild data with different image styles such as cartoon. Our code is available at: https://github.com/justacar/Plane-DUSt3R

三维重建布局估计预训练模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。