首个可生成全球尺度3D场景的生成模型,突破了传统AI的空间限制。
MetaEarth3D: Unlocking World-scale 3D Generation with Spatially Scalable Generative Modeling

- 基于1000万张全球分布的真实图像,构建空间可扩展的生成架构
- 支持跨千公里尺度的统一3D场景生成,兼具视觉与地理统计真实感
- 适用于地球观测、虚拟环境构建等超大空间智能任务
近期生成式AI在语言和视觉理解上取得显著进展,但其生成内容的空间尺度仍局限于有限环境,难以捕捉数千公里级地理环境的演变或大规模物理世界的结构。这一局限性制约了地球观测与仿真中的超大范围空间智能发展,暴露出生成式AI的一个深层差距:进展主要依赖模型参数与训练数据的规模扩张,而忽视了空间尺度作为智能的核心维度。为此,我们提出将空间尺度作为基础模型的新扩展轴,构建了首个具备全球尺度空间一致性的生成基础模型——MetaEarth3D。以光学地球观测模拟为测试场景,该模型可生成涵盖大范围地形、中等城市及精细街块的多层次、无界且多样化的3D场景。基于1000万张全球分布的真实世界训练图像,MetaEarth3D在视觉真实性和地理空间统计真实性方面表现优异。除生成能力外,它还可作为超大空间智能中多样化虚拟环境的生成数据引擎。本研究或有助于推动下一代地球观测空间智能的发展。
原文摘要 · Abstract (English)
Recent generative AI models have achieved remarkable breakthroughs in language and visual understanding. However, although these models can generate realistic visual content, their spatial scale remains confined to bounded environments, preventing them from capturing how geographic environments evolve across thousands of kilometers or from modeling the spatial structure of the large-scale physical world. This limitation poses a critical challenge for ultra-wide-area spatial intelligence in Earth observation and simulation, revealing a deeper gap in generative AI: progress has relied primarily on scaling model parameters and training data, while overlooking spatial scale as a core dimension of intelligence. Here, motivated by this missing dimension, we investigate spatial scale as a new scaling axis in foundation models and present MetaEarth3D, the first generative foundation model capable of spatially consistent generation at the planetary scale. Taking optical Earth observation simulation as a testbed, MetaEarth3D enables the generation of multi-level, unbounded, and diverse 3D scenes spanning large-scale terrains, medium-scale cities, and fine-grained street blocks. Built upon 10 million globally distributed real-world training images, MetaEarth3D demonstrates both strong visual realism and geospatial statistical realism. Beyond generation, MetaEarth3D serves as a generative data engine for diverse virtual environments in ultra-wide spatial intelligence. We argue that this study may help empower next-generation spatial intelligence for Earth observation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。