arXiv:2509.13414cs.CVcs.AI2025-09被引 317

一输入多任务的3D重建模型,直接输出精确场景几何与相机参数。

MapAnything: Universal Feed-Forward Metric 3D Reconstruction

  • 基于统一变压器架构,输入图像和几何信息后直接回归3D结构。
  • 在多个任务上超越或媲美专用模型,且训练更高效。
  • 适合需要统一3D重建框架的研究者与工业应用。

我们提出MapAnything,一种基于Transformer的统一前馈模型,可接收一张或多张图像及可选几何输入(如相机内参、位姿、深度或部分重建结果),并直接回归出度量空间下的3D场景几何与相机参数。该模型采用多视图场景几何的分解表示——包括深度图、局部射线图、相机位姿与度量尺度因子,有效将局部重建升级为全局一致的度量坐标系。通过标准化跨数据集的监督与训练策略,并支持灵活输入增强,MapAnything可在单次前馈中完成多种3D视觉任务,涵盖非标定结构光恢复、标定多视图立体匹配、单目深度估计、相机定位、深度补全等。大量实验与模型消融表明,其性能优于或媲美专业前馈模型,同时具备更高效的联合训练能力,为构建通用3D重建主干网络铺平道路。

原文摘要 · Abstract (English)

We introduce MapAnything, a unified transformer-based feed-forward model that ingests one or more images along with optional geometric inputs such as camera intrinsics, poses, depth, or partial reconstructions, and then directly regresses the metric 3D scene geometry and cameras. MapAnything leverages a factored representation of multi-view scene geometry, i.e., a collection of depth maps, local ray maps, camera poses, and a metric scale factor that effectively upgrades local reconstructions into a globally consistent metric frame. Standardizing the supervision and training across diverse datasets, along with flexible input augmentation, enables MapAnything to address a broad range of 3D vision tasks in a single feed-forward pass, including uncalibrated structure-from-motion, calibrated multi-view stereo, monocular depth estimation, camera localization, depth completion, and more. We provide extensive experimental analyses and model ablations demonstrating that MapAnything outperforms or matches specialist feed-forward models while offering more efficient joint training behavior, thus paving the way toward a universal 3D reconstruction backbone.

3D重建前馈模型统一框架深度估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。