从一张图生成可探索的3D虚拟世界,只对用户关注区域精细渲染。
NeoWorld: Neural Simulation of Explorable Virtual Worlds via Progressive 3D Unfolding
- 仅对用户探索区域生成3D细节,其余用2D合成,兼顾效率与真实感。
- 在WorldScore基准上显著优于现有2D和2.5D方法,支持自然语言控制物体行为。
- 适合游戏开发、数字孪生及沉浸式交互应用,尤其关注动态场景生成。
我们提出NeoWorld,一种基于深度学习的框架,可从单张输入图像生成可交互的3D虚拟世界。受1964年科幻小说《Simulacron-3》中按需构建世界的概念启发,该系统构建了广阔环境,仅对用户主动探索的区域以高视觉真实感呈现,采用以物体为中心的3D表示。不同于依赖全局生成或2D幻觉的以往方法,NeoWorld对关键前景物体实现完整3D建模,同时以2D方式合成背景和未交互区域,确保效率。这一混合场景结构结合前沿表示学习与物体到3D技术,支持灵活视角操作和物理合理的场景动画,允许用户通过自然语言指令控制物体外观与动态。随着用户交互,虚拟世界逐步展开并增加3D细节,提供动态、沉浸且视觉连贯的探索体验。NeoWorld在WorldScore基准上显著优于现有2D与深度分层的2.5D方法。
原文摘要 · Abstract (English)
We introduce NeoWorld, a deep learning framework for generating interactive 3D virtual worlds from a single input image. Inspired by the on-demand worldbuilding concept in the science fiction novel Simulacron-3 (1964), our system constructs expansive environments where only the regions actively explored by the user are rendered with high visual realism through object-centric 3D representations. Unlike previous approaches that rely on global world generation or 2D hallucination, NeoWorld models key foreground objects in full 3D, while synthesizing backgrounds and non-interacted regions in 2D to ensure efficiency. This hybrid scene structure, implemented with cutting-edge representation learning and object-to-3D techniques, enables flexible viewpoint manipulation and physically plausible scene animation, allowing users to control object appearance and dynamics using natural language commands. As users interact with the environment, the virtual world progressively unfolds with increasing 3D detail, delivering a dynamic, immersive, and visually coherent exploration experience. NeoWorld significantly outperforms existing 2D and depth-layered 2.5D methods on the WorldScore benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。