arXiv:2506.08006cs.CV2025-06被引 5

用物理模拟器+生成模型,实现可控且逼真的世界生成。

Dreamland: Controllable World Creation with Simulator and Generative Models

  • 分层抽象中间表示,打通模拟器与生成模型的语义鸿沟。
  • 图像质量提升50.8%,控制能力增强17.9%,优于现有方法。
  • 适合需要高可控性场景编辑与智能体训练的研究者使用。

大规模视频生成模型可合成多样且逼真的视觉内容以构建动态世界,但往往缺乏逐元素可控性,限制其在场景编辑和具身智能体训练中的应用。我们提出Dreamland,一种融合物理仿真器粒度控制与大规模预训练生成模型逼真输出的混合世界生成框架。特别地,设计了一种分层世界抽象结构,将像素级与对象级语义及几何信息编码为中间表示,以连接仿真器与生成模型。该方法提升了可控性,通过早期对齐真实世界分布降低了适配成本,并支持即插即用现有及未来预训练生成模型。我们进一步构建了D3Sim数据集,以促进混合生成流水线的训练与评估。实验表明,Dreamland在图像质量上较基线提升50.8%,可控性增强17.9%,在增强具身智能体训练方面具有巨大潜力。代码与数据将公开。

原文摘要 · Abstract (English)

Large-scale video generative models can synthesize diverse and realistic visual content for dynamic world creation, but they often lack element-wise controllability, hindering their use in editing scenes and training embodied AI agents. We propose Dreamland, a hybrid world generation framework combining the granular control of a physics-based simulator and the photorealistic content output of large-scale pretrained generative models. In particular, we design a layered world abstraction that encodes both pixel-level and object-level semantics and geometry as an intermediate representation to bridge the simulator and the generative model. This approach enhances controllability, minimizes adaptation cost through early alignment with real-world distributions, and supports off-the-shelf use of existing and future pretrained generative models. We further construct a D3Sim dataset to facilitate the training and evaluation of hybrid generation pipelines. Experiments demonstrate that Dreamland outperforms existing baselines with 50.8% improved image quality, 17.9% stronger controllability, and has great potential to enhance embodied agent training. Code and data will be made available.

世界生成可控生成具身智能物理模拟

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。