arXiv:2601.11514cs.CVcs.LG2026-01被引 14

从随意拍摄视频生成精准3D模型,应对遮挡与杂乱背景挑战。

ShapeR: Robust Conditional 3D Shape Generation from Casual Captures

  • 融合视觉惯性定位、多视角图像和机器描述,联合建模物体外观与空间关系。
  • 在真实场景数据上,相比最强方法,重建误差降低2.7倍(Chamfer距离)。
  • 适合需要从手机拍摄视频生成高质量3D模型的开发者或设计师使用。

近期3D形状生成取得显著进展,但多数方法依赖干净、无遮挡且分割良好的输入,这在真实场景中罕见。本文提出ShapeR,一种从随意拍摄序列生成条件3D物体形状的新方法。给定图像序列,利用现成的视觉惯性SLAM、3D检测算法和视觉语言模型,为每个物体提取稀疏SLAM点、带姿态的多视角图像及机器生成的描述文本。通过一个经过校正的流变换器,有效融合这些模态信息,生成高保真度的度量3D形状。为提升对随意采集数据的鲁棒性,采用实时组合增强、跨物体与场景层级的数据集课程训练策略,并设计处理背景杂乱的方法。此外,构建了一个新评估基准,包含7个真实场景中的178个野外物体,附带几何标注。实验表明,ShapeR在此严苛设置下显著优于现有方法,相较于最先进水平,Chamfer距离提升2.7倍。

原文摘要 · Abstract (English)

Recent advances in 3D shape generation have achieved impressive results, but most existing methods rely on clean, unoccluded, and well-segmented inputs. Such conditions are rarely met in real-world scenarios. We present ShapeR, a novel approach for conditional 3D object shape generation from casually captured sequences. Given an image sequence, we leverage off-the-shelf visual-inertial SLAM, 3D detection algorithms, and vision-language models to extract, for each object, a set of sparse SLAM points, posed multi-view images, and machine-generated captions. A rectified flow transformer trained to effectively condition on these modalities then generates high-fidelity metric 3D shapes. To ensure robustness to the challenges of casually captured data, we employ a range of techniques including on-the-fly compositional augmentations, a curriculum training scheme spanning object- and scene-level datasets, and strategies to handle background clutter. Additionally, we introduce a new evaluation benchmark comprising 178 in-the-wild objects across 7 real-world scenes with geometry annotations. Experiments show that ShapeR significantly outperforms existing approaches in this challenging setting, achieving an improvement of 2.7x in Chamfer distance compared to state of the art.

3D生成视觉定位多模态真实场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。