将真实场景转为可编辑的Minecraft世界,支持智能体训练
World2Minecraft: Occupancy-Driven Simulated Scenes Construction

- 基于3D语义占据预测,自动构建结构化游戏场景
- 打造10万+图像的大规模占据数据集,显著提升重建质量
- 适合做具身智能、视觉语言导航研究的个性化实验平台
具身智能需要高保真仿真环境以支持感知与决策,但现有平台常存在数据污染和灵活性不足的问题。为此,我们提出World2Minecraft,通过3D语义占据预测将真实场景转换为结构化的Minecraft环境。重建后的场景可轻松用于视觉-语言导航(VLN)等下游任务。然而,我们发现重建质量高度依赖准确的占据预测,而现有模型受限于数据稀缺与泛化能力差。为此,我们设计了一种低成本、自动化且可扩展的数据采集流程,构建了MinecraftOcc数据集,包含来自156个丰富室内场景的100,165张图像。大量实验表明,该数据集对现有数据形成有效补充,并对当前主流方法构成显著挑战。这些成果推动了占据预测性能提升,凸显World2Minecraft在个性化具身智能研究中的可定制与可编辑价值。
原文摘要 · Abstract (English)
Embodied intelligence requires high-fidelity simulation environments to support perception and decision-making, yet existing platforms often suffer from data contamination and limited flexibility. To mitigate this, we propose World2Minecraft to convert real-world scenes into structured Minecraft environments based on 3D semantic occupancy prediction. In the reconstructed scenes, we can effortlessly perform downstream tasks such as Vision-Language Navigation(VLN). However, we observe that reconstruction quality heavily depends on accurate occupancy prediction, which remains limited by data scarcity and poor generalization in existing models. We introduce a low-cost, automated, and scalable data acquisition pipeline for creating customized occupancy datasets, and demonstrate its effectiveness through MinecraftOcc, a large-scale dataset featuring 100,165 images from 156 richly detailed indoor scenes. Extensive experiments show that our dataset provides a critical complement to existing datasets and poses a significant challenge to current SOTA methods. These findings contribute to improving occupancy prediction and highlight the value of World2Minecraft in providing a customizable and editable platform for personalized embodied AI research. Project page:https://world2minecraft.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。