用多视角视频生成上百个物体的组合3D场景,解决遮挡难题。
WorldSculpt: Generating Compositional Worlds from Grounded Videos

- 基于单物体生成模型,通过多视角条件控制实现复杂场景组合生成
- 在遮挡严重场景中仍保持高精度,优于现有方法且随复杂度提升优势更大
- 适用于游戏、VR/AR等下游应用,可将已有3DGS世界转为网格场景
我们研究如何生成包含数百个物体的稠密杂乱场景的组合式3D表示。目标是将场景表示为一组位于共享世界坐标系中的独立物体网格,满足游戏、AR/VR、仿真和机器人等下游应用需求。该任务在密集遮挡场景中极具挑战,因物体相互遮挡,每个视角仅能观测到部分几何信息。基于几何的方法通常重建为单一整体表示,导致遮挡区域几何不完整;而现有具有生成先验的组合方法主要局限于较简单场景。本文表明,可通过将强大的单物体3D生成先验适配多视角观测,实现复杂场景的组合式生成。我们以Pixal3D为例,引入多视角条件路径,使物体生成基于多个姿态观测进行定位。尽管模型仅在规范空间的单物体上微调,却能在无场景级训练的情况下泛化至严重遮挡的大场景,验证了该范式的可行性与可扩展性。我们进一步构建了UE-MeshyScene——一个包含数百个物体、每物体标注及真实网格的逼真稠密场景基准数据集。在单物体、受控多物体和UE-MeshyScene评估中,本方法始终优于先前方法,且随着场景复杂度与遮挡程度增加,性能提升更显著。最后,我们展示了广泛适用性:将生成的3DGS世界(如Marble和HY-World 2.0)转换为组合网格场景。
原文摘要 · Abstract (English)
We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds of objects. The goal is to represent the scene as a collection of individual object meshes placed in a shared world frame, as required by downstream applications such as gaming, AR/VR, simulation, and robotics. This task is challenging in densely cluttered scenes, where objects heavily occlude one another and each view reveals only a fraction of their geometry. Geometry-based approaches typically reconstruct the scene as a single representation and leave incomplete geometry in occluded regions, while existing compositional methods with generative priors are largely limited to relatively simple scenes. We show that complex scenes with hundreds of objects can instead be generated compositionally by adapting a strong single-object 3D generative prior to multi-view observations. We instantiate this paradigm with Pixal3D, extending it with a multi-view conditioning pathway that grounds object generation in multiple posed observations. Although the model is finetuned entirely on single objects in canonical space, it generalizes to large scenes with severe occlusion without any scene-level training, demonstrating the feasibility and scalability of this paradigm. We further introduce UE-MeshyScene, a photorealistic benchmark of densely cluttered scenes with hundreds of objects, per-object annotations, and ground-truth meshes. Across single-object, controlled multi-object, and UE-MeshyScene evaluations, our method consistently outperforms prior approaches, with larger gains as scene complexity and occlusion increase. Finally, we demonstrate broader applicability by converting generated 3DGS worlds, such as Marble and HY-World 2.0, into compositional mesh scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。