构建4万场景的高真实感室内数据集,解决布局失真与物体碰撞问题。
InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts
- 融合真实扫描、程序生成与设计场景,生成4万多样室内场景。
- 平均每区域41.5个物体,含196万3D物件,覆盖288类物品与15种空间类型。
- 支持物理仿真去碰撞,适配智能体导航与场景生成任务研究。
Embodied AI的发展依赖于大规模、可模拟的3D场景数据集,需具备场景多样性与真实布局。现有数据集普遍存在规模有限、布局简化(缺少小物件)、严重物体重叠等问题。为此,我们提出【InternScenes】——一个大规模可模拟的室内场景数据集,整合真实扫描、程序生成与设计师创建的场景,共包含约40,000个多样化场景,涵盖1.96M个3D对象,覆盖15种常见场景类型与288个物体类别。特别保留大量小型物品,实现平均每区域41.5个物体的真实复杂布局。通过完整数据处理流程,为真实扫描创建“现实-仿真”复制品,增强场景可交互性,并利用物理模拟消除物体碰撞。我们在两个基准任务中验证其价值:场景布局生成与点目标导航,均揭示了复杂真实布局带来的新挑战。更重要的是,InternScenes为这两项任务的大规模模型训练提供可能,使复杂场景下的生成与导航成为现实。我们将开源数据、模型与基准测试,以促进社区发展。
原文摘要 · Abstract (English)
The advancement of Embodied AI heavily relies on large-scale, simulatable 3D scene datasets characterized by scene diversity and realistic layouts. However, existing datasets typically suffer from limitations in data scale or diversity, sanitized layouts lacking small items, and severe object collisions. To address these shortcomings, we introduce \textbf{InternScenes}, a novel large-scale simulatable indoor scene dataset comprising approximately 40,000 diverse scenes by integrating three disparate scene sources, real-world scans, procedurally generated scenes, and designer-created scenes, including 1.96M 3D objects and covering 15 common scene types and 288 object classes. We particularly preserve massive small items in the scenes, resulting in realistic and complex layouts with an average of 41.5 objects per region. Our comprehensive data processing pipeline ensures simulatability by creating real-to-sim replicas for real-world scans, enhances interactivity by incorporating interactive objects into these scenes, and resolves object collisions by physical simulations. We demonstrate the value of InternScenes with two benchmark applications: scene layout generation and point-goal navigation. Both show the new challenges posed by the complex and realistic layouts. More importantly, InternScenes paves the way for scaling up the model training for both tasks, making the generation and navigation in such complex scenes possible. We commit to open-sourcing the data, models, and benchmarks to benefit the whole community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。