arXiv:2601.19048cs.CV2026-01被引 1

用少量图像生成可控大场景,效率更高。

NuiWorld: Exploring a Scalable Framework for End-to-End Controllable World Generation

  • 用伪草图控制场景生成,支持新草图泛化。
  • 通过生成式数据扩充解决数据稀缺问题。
  • 分块压缩表示降低计算量,适合大场景建模。

世界生成是视频游戏、仿真和机器人应用的核心能力。然而现有方法面临可控性、可扩展性和效率三大挑战:端到端场景生成受数据稀缺限制;基于物体中心的生成依赖固定分辨率,大场景保真度下降;训练无关方法虽灵活但推理慢且耗能高。我们提出NuiWorld框架,通过生成式自举策略,从少数输入图像出发,结合近期3D重建与可扩展场景生成技术,合成多尺寸、多布局的场景数据,为端到端模型训练提供充足数据。通过伪草图标签实现可控生成,并展示对未见草图的一定泛化能力。将场景表示为可变大小的场景块集合,压缩为扁平向量集,显著减少大场景的标记长度,在保持几何保真度的同时提升训练与推理效率。

原文摘要 · Abstract (English)

World generation is a fundamental capability for applications like video games, simulation, and robotics. However, existing approaches face three main obstacles: controllability, scalability, and efficiency. End-to-end scene generation models have been limited by data scarcity. While object-centric generation approaches rely on fixed resolution representations, degrading fidelity for larger scenes. Training-free approaches, while flexible, are often slow and computationally expensive at inference time. We present NuiWorld, a framework that attempts to address these challenges. To overcome data scarcity, we propose a generative bootstrapping strategy that starts from a few input images. Leveraging recent 3D reconstruction and expandable scene generation techniques, we synthesize scenes of varying sizes and layouts, producing enough data to train an end-to-end model. Furthermore, our framework enables controllability through pseudo sketch labels, and demonstrates a degree of generalization to previously unseen sketches. Our approach represents scenes as a collection of variable scene chunks, which are compressed into a flattened vector-set representation. This significantly reduces the token length for large scenes, enabling consistent geometric fidelity across scenes sizes while improving training and inference efficiency.

世界生成可控生成3D建模高效建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。