仅用一张图就能生成带纹理和布局的3D模型,适合复杂场景。
SAM 3D: 3Dfy Anything in Images

- 通过人机协作标注海量3D数据,构建高质量训练集。
- 在真实场景中人类偏好测试胜率超5:1,显著优于现有方法。
- 适合做3D重建、虚拟现实或工业设计的研究者与开发者。
我们提出SAM 3D,一种基于视觉的3D物体重建生成模型,可从单张图像中预测几何、纹理和布局。该模型在存在遮挡和场景杂乱的自然图像中表现优异,得益于上下文视觉线索的有效利用。我们采用人机协同流程标注物体形状、纹理和姿态,构建了前所未有的大规模视觉对齐3D重建数据集。通过融合合成预训练与真实世界对齐的多阶段训练框架,突破了3D数据瓶颈。在真实物体与场景上的人类偏好测试中,取得至少5:1的胜率优势。我们将开源代码、模型权重、在线演示及一个新的野外3D重建挑战基准。
原文摘要 · Abstract (English)
We present SAM 3D, a generative model for visually grounded 3D object reconstruction, predicting geometry, texture, and layout from a single image. SAM 3D excels in natural images, where occlusion and scene clutter are common and visual recognition cues from context play a larger role. We achieve this with a human- and model-in-the-loop pipeline for annotating object shape, texture, and pose, providing visually grounded 3D reconstruction data at unprecedented scale. We learn from this data in a modern, multi-stage training framework that combines synthetic pretraining with real-world alignment, breaking the 3D "data barrier". We obtain significant gains over recent work, with at least a 5:1 win rate in human preference tests on real-world objects and scenes. We will release our code and model weights, an online demo, and a new challenging benchmark for in-the-wild 3D object reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。