用智能代理生成可自由探索的大规模3D世界,支持编辑复用。
WorldClaw: Agentic 3D Open-World Generation at Scale

- 分阶段生成:先规划结构,再细化细节,确保全局一致。
- 生成的3D场景空间连贯、内容丰富,支持逐个物体编辑。
- 适合游戏开发、虚拟现实等需要可交互3D环境的场景。
从开放文本生成大规模、可自由探索的3D世界仍具挑战,因系统需同时保持全局空间一致性、丰富的局部内容,并生成适用于下游编辑与复用的显式资产。我们提出WorldClaw,一种全代理式、粗到细的开放世界3D场景生成框架。规划代理将文本提示转化为区域、地形、资产、材质及空间关系的结构化说明。WorldClaw随后基于语义布局、可复用资产、生成或程序化材质以及区域感知高度图构建全局一致的地形基础。对于细节需求高的区域,它生成地形条件下的组合内容,重建可编辑的纹理网格并恢复其在地形上的位置;渲染驱动的代理进一步优化地形、物体、外观与接触关系。在多种开放世界提示下,WorldClaw生成了具有连贯空间组织、视觉吸引人的局部内容和可编辑实例级资产的大规模场景,同时保持一致的全局地形结构。
原文摘要 · Abstract (English)
Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jointly maintain global spatial coherence, rich local content, and explicit assets suitable for downstream editing and reuse. We present WorldClaw, a fully agentic, coarse-to-fine framework for open-world 3D scene generation. Planning agents translate a text prompt into a structured specification of regions, terrain, assets, materials, and spatial relations. WorldClaw then builds a globally coherent terrain foundation from semantic layouts, reusable assets, generative or procedural materials, and a region-aware height field. For detail-demanding regions, it generates terrain-conditioned compositions, reconstructs editable textured meshes, and recovers their placement on the terrain; render-based agents further refine terrain, objects, appearance, and contacts. Across diverse open-world prompts, WorldClaw produces large-scale scenes with coherent spatial organization, visually compelling local content, and editable instance-level assets while preserving a consistent global terrain structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。