arXiv:2605.00781cs.CV2026-05

用任意形状地图生成3D世界,解决尺度不一致问题

Map2World: Segment Map Conditioned Text to 3D World Generation

论文配图:Map2World: Segment Map Conditioned Text to 3D World Generation
图 1 · 摘自论文原文
  • 根据用户自定义的分段地图生成3D场景,支持复杂布局
  • 引入细节增强网络,提升物体细节且保持整体一致性
  • 适用于沉浸式内容创作与自动驾驶仿真,可控性强

3D世界生成在沉浸式内容创作和自动驾驶仿真中至关重要。现有方法受限于网格布局,普遍存在物体尺度不一致的问题。本文提出Map2World框架,首次实现基于用户自定义任意形状与尺度分段地图的3D世界生成,确保大范围环境下的全局尺度一致性与灵活性。为提升生成质量,设计细节增强网络,在不破坏场景整体连贯性的前提下融入全局结构信息,生成精细细节。整个流程利用资产生成器中的强先验知识,即使在有限训练数据下也能实现跨领域的鲁棒泛化。大量实验表明,本方法在用户可控性、尺度一致性与内容连贯性上显著优于现有方法,支持更复杂条件下的3D世界生成。

原文摘要 · Abstract (English)

3D world generation is essential for applications such as immersive content creation or autonomous driving simulation. Recent advances in 3D world generation have shown promising results; however, these methods are constrained by grid layouts and suffer from inconsistencies in object scale throughout the entire world. In this work, we introduce a novel framework, Map2World, that first enables 3D world generation conditioned on user-defined segment maps of arbitrary shapes and scales, ensuring global-scale consistency and flexibility across expansive environments. To further enhance the quality, we propose a detail enhancer network that generates fine details of the world. The detail enhancer enables the addition of fine-grained details without compromising overall scene coherence by incorporating global structure information. We design the entire pipeline to leverage strong priors from asset generators, achieving robust generalization across diverse domains, even under limited training data for scene generation. Extensive experiments demonstrate that our method significantly outperforms existing approaches in user-controllability, scale consistency, and content coherence, enabling users to generate 3D worlds under more complex conditions.

3D生成地图生成场景一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。