arXiv:2410.05514cs.CVcs.AI2024-10CoRL被引 6

用3D扩散模型实现多类别物体在稀疏视角下的精准三维重建。

Toward General Object-level Mapping from Sparse Views with 3D Diffusion Priors

  • 引入预训练3D扩散模型作为形状先验,支持多类别物体
  • 在真实场景数据集上优于当前最佳方法,稀疏视图下重建更准确
  • 无需微调模型,通过非线性约束融合传感器数据与生成先验

物体级映射旨在从多视角传感器观测中构建场景内物体的详细三维形状与位姿。传统方法因遮挡和传感器噪声难以完整重建形状并准确估计位姿,且需密集观测覆盖所有物体,这在机器人轨迹中难以实现。近期工作虽引入生成形状先验以支持稀疏视图,但仅限于单类别物体。本文提出通用物体级映射系统GOM,利用3D扩散模型作为多类别形状先验,输出所有物体的神经辐射场(NeRF)以同时表示纹理与几何。GOM采用有效公式,在不微调预训练扩散模型的前提下,通过传感器测量施加额外非线性约束进行引导。我们还设计了一种概率优化框架,联合融合多视角观测与扩散先验,实现物体位姿与形状的联合估计。GOM在真实世界基准测试中展现出优异的多类别映射性能,相比现有最优方法在稀疏视图下获得更精确的重建结果。代码将公开:https://github.com/TRAILab/GeneralObjectMapping。

原文摘要 · Abstract (English)

Object-level mapping builds a 3D map of objects in a scene with detailed shapes and poses from multi-view sensor observations. Conventional methods struggle to build complete shapes and estimate accurate poses due to partial occlusions and sensor noise. They require dense observations to cover all objects, which is challenging to achieve in robotics trajectories. Recent work introduces generative shape priors for object-level mapping from sparse views, but is limited to single-category objects. In this work, we propose a General Object-level Mapping system, GOM, which leverages a 3D diffusion model as shape prior with multi-category support and outputs Neural Radiance Fields (NeRFs) for both texture and geometry for all objects in a scene. GOM includes an effective formulation to guide a pre-trained diffusion model with extra nonlinear constraints from sensor measurements without finetuning. We also develop a probabilistic optimization formulation to fuse multi-view sensor observations and diffusion priors for joint 3D object pose and shape estimation. Our GOM system demonstrates superior multi-category mapping performance from sparse views, and achieves more accurate mapping results compared to state-of-the-art methods on the real-world benchmarks. We will release our code: https://github.com/TRAILab/GeneralObjectMapping.

3D重建扩散模型多类别稀疏视图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。