arXiv:2507.21809cs.CV2025-07被引 94

用文字或图片生成可探索互动的3D世界,支持全景体验和游戏级兼容。

HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels

  • 用全景图像做世界代理,分层构建语义3D网格。
  • 生成结果支持导出网格,适配现有图形管线。
  • 对象解耦表示,提升交互灵活性,适合虚拟现实开发。

从文本或图像生成沉浸式、可探索且可交互的3D世界仍是计算机视觉与图形学中的核心挑战。现有方法主要分为两类:基于视频的方法虽多样性丰富但缺乏3D一致性与渲染效率;基于3D的方法虽几何一致但受限于训练数据少和内存效率低。为此,我们提出HunyuanWorld 1.0,融合两者优势,实现从文本或图像生成沉浸式、可探索、可交互的3D场景。其核心特点包括:1)通过全景世界代理实现360°沉浸体验;2)支持网格导出,无缝对接现有图形管线;3)解耦对象表示,增强交互性。框架采用语义分层的3D网格表示,以全景图像为360°世界代理,实现语义感知的世界分解与重建,从而生成多样化3D世界。大量实验表明,该方法在生成连贯、可探索、可交互的3D世界方面达到当前最优性能,并在虚拟现实、物理仿真、游戏开发与互动内容创作中展现广泛应用潜力。

原文摘要 · Abstract (English)

Creating immersive and playable 3D worlds from texts or images remains a fundamental challenge in computer vision and graphics. Existing world generation approaches typically fall into two categories: video-based methods that offer rich diversity but lack 3D consistency and rendering efficiency, and 3D-based methods that provide geometric consistency but struggle with limited training data and memory-inefficient representations. To address these limitations, we present HunyuanWorld 1.0, a novel framework that combines the best of both worlds for generating immersive, explorable, and interactive 3D scenes from text and image conditions. Our approach features three key advantages: 1) 360° immersive experiences via panoramic world proxies; 2) mesh export capabilities for seamless compatibility with existing computer graphics pipelines; 3) disentangled object representations for augmented interactivity. The core of our framework is a semantically layered 3D mesh representation that leverages panoramic images as 360° world proxies for semantic-aware world decomposition and reconstruction, enabling the generation of diverse 3D worlds. Extensive experiments demonstrate that our method achieves state-of-the-art performance in generating coherent, explorable, and interactive 3D worlds while enabling versatile applications in virtual reality, physical simulation, game development, and interactive content creation.

3D生成虚拟现实交互设计全景建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。