用轻量代理与纹理生成逼真沉浸式世界,支持移动端实时渲染。
ImmerseGen: Agent-Guided Immersive World Generation with Alpha-Textured Proxies
- 用轻量几何代理+合成贴图构建场景,解耦真实感与复杂建模。
- 通过地形与上下文感知贴图,生成视觉连贯的多样化世界。
- 基于视觉语言模型的智能体实现精准布物与多模态增强。
自动化沉浸式虚拟现实场景生成仍是主要研究挑战。现有方法通常依赖复杂几何结构并进行后期简化,导致流程低效或真实感有限。本文提出 ImmerseGen,一种新型代理引导的紧凑且逼真的世界生成框架,将真实感与详尽几何建模解耦。ImmerseGen 将场景表示为分层轻量级几何代理及其合成的 RGBA 贴图,支持移动端 VR 头显的实时渲染。我们提出基于地形的贴图生成方法,结合情境感知贴图用于景物构建,实现多样且视觉一致的世界生成。基于视觉语言模型(VLM)的智能体利用语义网格分析实现精准资产布局,并通过视觉动态和环境音等多模态增强丰富场景。实验与实时 VR 应用表明,ImmerseGen 在保真度、空间连贯性与渲染效率方面均优于现有方法。
原文摘要 · Abstract (English)
Automating immersive VR scene creation remains a primary research challenge. Existing methods typically rely on complex geometry with post-simplification, resulting in inefficient pipelines or limited realism. In this paper, we introduce ImmerseGen, a novel agent-guided framework for compact and photorealistic world generation that decouples realism from exhaustive geometric modeling. ImmerseGen represents scenes as hierarchical compositions of lightweight geometric proxies with synthesized RGBA textures, facilitating real-time rendering on mobile VR headsets. We propose terrain-conditioned texturing for base world generation, combined with context-aware texturing for scenery, to produce diverse and visually coherent worlds. VLM-based agents employ semantic grid-based analysis for precise asset placement and enrich scenes with multimodal enhancements such as visual dynamics and ambient sound. Experiments and real-time VR applications demonstrate that ImmerseGen achieves superior photorealism, spatial coherence, and rendering efficiency compared to existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。