arXiv:2503.17316cs.CV2025-03CVPR被引 70

Pow3r利用相机与场景先验,实现多模态3D重建的高精度推理。

Pow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene Priors

  • 统一网络融合图像、相机参数、深度等多源信息作为条件输入
  • 在多个任务上达到当前最优,支持原图分辨率推理与点云补全
  • 适合需要灵活使用先验信息的3D视觉应用开发者

我们提出Pow3r,一种新型大规模3D视觉回归模型,可灵活接收多种输入模态。不同于以往前馈模型无法在测试时利用已知相机或场景先验,Pow3r可在单一网络中整合任意组合的辅助信息,如内参、相对位姿、稠密或稀疏深度以及输入图像。基于近期的DUSt3R范式——一种利用强大预训练的Transformer架构,我们的轻量级、多功能条件化机制为网络提供额外引导,在有辅助信息时提升预测精度。训练时每轮随机输入不同模态子集,使模型能在测试时适应不同先验水平。这带来了新能力,如原图分辨率推理和点云补全。在3D重建、深度补全、多视角深度预测、多视图立体匹配和多视图位姿估计任务上,实验结果均达到最先进水平,验证了Pow3r有效利用所有可用信息的能力。项目主页:https://europe.naverlabs.com/pow3r。

原文摘要 · Abstract (English)

We present Pow3r, a novel large 3D vision regression model that is highly versatile in the input modalities it accepts. Unlike previous feed-forward models that lack any mechanism to exploit known camera or scene priors at test time, Pow3r incorporates any combination of auxiliary information such as intrinsics, relative pose, dense or sparse depth, alongside input images, within a single network. Building upon the recent DUSt3R paradigm, a transformer-based architecture that leverages powerful pre-training, our lightweight and versatile conditioning acts as additional guidance for the network to predict more accurate estimates when auxiliary information is available. During training we feed the model with random subsets of modalities at each iteration, which enables the model to operate under different levels of known priors at test time. This in turn opens up new capabilities, such as performing inference in native image resolution, or point-cloud completion. Our experiments on 3D reconstruction, depth completion, multi-view depth prediction, multi-view stereo, and multi-view pose estimation tasks yield state-of-the-art results and confirm the effectiveness of Pow3r at exploiting all available information. The project webpage is https://europe.naverlabs.com/pow3r.

3D重建多模态深度补全视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。