arXiv:2502.07685cs.CV2025-02CVPR被引 42

一个模型搞定三维重建全流程,还能精细控制生成过程。

Matrix3D: Large Photogrammetry Model All-in-One

  • 用多模态扩散变换器统一处理图像、相机参数和深度图
  • 在部分数据条件下仍可全模态训练,大幅提升可用数据量
  • 支持多轮交互控制,适合3D内容创作与高精度重建

我们提出Matrix3D,一个统一模型,仅用同一架构即可完成姿态估计、深度预测和新视角合成等摄影测量子任务。该模型采用多模态扩散变换器(DiT),融合图像、相机参数与深度图等多种模态信息。其大规模多模态训练的关键在于引入掩码学习策略,即使在部分完整数据(如图像-姿态对、图像-深度对)条件下,也能实现全模态训练,显著扩充可用训练数据规模。Matrix3D在姿态估计和新视角合成任务上达到当前最优性能,并通过多轮交互提供细粒度控制能力,是3D内容创作的创新工具。项目主页:https://nju-3dv.github.io/projects/matrix3d。

原文摘要 · Abstract (English)

We present Matrix3D, a unified model that performs several photogrammetry subtasks, including pose estimation, depth prediction, and novel view synthesis using just the same model. Matrix3D utilizes a multi-modal diffusion transformer (DiT) to integrate transformations across several modalities, such as images, camera parameters, and depth maps. The key to Matrix3D's large-scale multi-modal training lies in the incorporation of a mask learning strategy. This enables full-modality model training even with partially complete data, such as bi-modality data of image-pose and image-depth pairs, thus significantly increases the pool of available training data. Matrix3D demonstrates state-of-the-art performance in pose estimation and novel view synthesis tasks. Additionally, it offers fine-grained control through multi-round interactions, making it an innovative tool for 3D content creation. Project page: https://nju-3dv.github.io/projects/matrix3d.

三维重建扩散模型多模态生成式建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。