arXiv:2512.13683cs.CV2025-12被引 3

用3D实例模型隐式学习空间关系,无需场景数据也能生成新布局。

I-Scene: 3D Instance Models are Implicit Generalizable Spatial Learners

  • 用模型自身空间监督替代数据集限制,让生成器学会空间推理。
  • 随机组合物体也能推断远近、支撑、对称等关系,准确率显著提升。
  • 适合做交互式3D场景生成的通用基础模型,尤其关注空间理解者。

泛化仍是交互式3D场景生成的核心挑战。现有方法依赖有限场景数据集进行空间理解,限制了对新布局的泛化能力。本文重新编程预训练的3D实例生成器,使其作为场景级学习者,用模型内生的空间监督取代数据依赖的监督信号。该重编程释放了生成器可迁移的空间知识,使其能泛化至未见布局与新颖物体组合。令人惊讶的是,即使训练场景由随机物体组成,空间推理仍能自发涌现。这表明生成器的可迁移场景先验,能从纯几何线索中有效推断邻近性、支撑关系与对称性。我们摒弃传统规范空间,提出以视角为中心的场景空间建模,实现全前馈、可泛化的场景生成器,直接从实例模型学习空间关系。定量与定性结果表明,3D实例生成器是隐式的空间学习者与推理者,为交互式3D场景理解与生成奠定了基础模型方向。

原文摘要 · Abstract (English)

Generalization remains the central challenge for interactive 3D scene generation. Existing learning-based approaches ground spatial understanding in limited scene dataset, restricting generalization to new layouts. We instead reprogram a pre-trained 3D instance generator to act as a scene level learner, replacing dataset-bounded supervision with model-centric spatial supervision. This reprogramming unlocks the generator transferable spatial knowledge, enabling generalization to unseen layouts and novel object compositions. Remarkably, spatial reasoning still emerges even when the training scenes are randomly composed objects. This demonstrates that the generator's transferable scene prior provides a rich learning signal for inferring proximity, support, and symmetry from purely geometric cues. Replacing widely used canonical space, we instantiate this insight with a view-centric formulation of the scene space, yielding a fully feed-forward, generalizable scene generator that learns spatial relations directly from the instance model. Quantitative and qualitative results show that a 3D instance generator is an implicit spatial learner and reasoner, pointing toward foundation models for interactive 3D scene understanding and generation. Project page: https://luling06.github.io/I-Scene-project/

3D生成空间推理基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。