用真实扫描自动构建可交互3D场景,降低人工成本。
MetaScenes: Towards Automated Replica Creation for Real-world 3D Scans
- 基于真实扫描自动替换物体,无需人工设计。
- 含1.5万+物体、831类,支持机器人操作与视觉语言导航。
- 适合做具身智能、仿真到现实迁移的研究者。
具身人工智能(EAI)研究需要高质量、多样化的3D场景以支持技能习得、仿真到现实的迁移及泛化能力。然而,达到这些质量标准需精确复现真实世界中的物体多样性,现有数据集高度依赖艺术家设计,人力成本高且难以扩展。为此,我们提出MetaScenes——一个大规模、可仿真的3D场景数据集,源自真实扫描,包含15,366个物体,覆盖831个细粒度类别。进一步提出Scan2Sim,一种鲁棒的多模态对齐模型,实现资产的自动化、高质量替换,彻底摆脱对人工设计的依赖。我们还构建了两个评估基准:针对机器人操作的小物品布局生成任务,以及视觉-语言导航(VLN)中的域迁移任务,验证跨域迁移能力。结果表明,MetaScenes能有效提升EAI中代理学习的泛化性与仿真到现实应用潜力,为该领域带来新可能。
原文摘要 · Abstract (English)
Embodied AI (EAI) research requires high-quality, diverse 3D scenes to effectively support skill acquisition, sim-to-real transfer, and generalization. Achieving these quality standards, however, necessitates the precise replication of real-world object diversity. Existing datasets demonstrate that this process heavily relies on artist-driven designs, which demand substantial human effort and present significant scalability challenges. To scalably produce realistic and interactive 3D scenes, we first present MetaScenes, a large-scale, simulatable 3D scene dataset constructed from real-world scans, which includes 15366 objects spanning 831 fine-grained categories. Then, we introduce Scan2Sim, a robust multi-modal alignment model, which enables the automated, high-quality replacement of assets, thereby eliminating the reliance on artist-driven designs for scaling 3D scenes. We further propose two benchmarks to evaluate MetaScenes: a detailed scene synthesis task focused on small item layouts for robotic manipulation and a domain transfer task in vision-and-language navigation (VLN) to validate cross-domain transfer. Results confirm MetaScene's potential to enhance EAI by supporting more generalizable agent learning and sim-to-real applications, introducing new possibilities for EAI research. Project website: https://meta-scenes.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。