构建大规模3D室内布局数据集,支持文本生成复杂场景。
M3DLayout: A Multi-Source Dataset of 3D Indoor Layouts and Structured Descriptions for 3D Generation
- 整合真实扫描、CAD设计与程序生成三类数据源
- 含21,367个布局和43.3万+物体实例,标注细致
- 适合研究文本驱动3D生成的学者与开发者使用
在文本驱动的3D场景生成中,物体布局作为连接高层语言指令与详细几何输出的关键中间表示,不仅提供物理合理的结构蓝图,还支持语义可控性和交互编辑。然而,现有3D室内布局生成模型的学习能力受限于数据集规模小、多样性不足和标注质量差。为此,我们提出M3DLayout——一个大规模、多源的3D室内布局生成数据集。该数据集包含21,367个布局和超过433,000个物体实例,融合真实扫描、专业CAD设计及程序生成场景三类来源。每个布局均配有结构化文本描述,涵盖全局场景摘要、大型家具的相对位置关系以及小型物品的精细布局。这一多样且丰富标注的数据资源,使模型能够学习跨多种室内环境的复杂空间与语义模式。为评估其潜力,我们基于文本条件扩散模型和自回归模型建立基准测试。实验表明,该数据集为布局生成模型训练提供了坚实基础。其多源特性显著提升了多样性,尤其由Inf3DLayout子集提供的丰富小物体信息,使生成场景更复杂、更细致。所有数据与代码将在论文接受后公开。
原文摘要 · Abstract (English)
In text-driven 3D scene generation, object layout serves as a crucial intermediate representation that bridges high-level language instructions with detailed geometric output. It not only provides a structural blueprint for ensuring physical plausibility but also supports semantic controllability and interactive editing. However, the learning capabilities of current 3D indoor layout generation models are constrained by the limited scale, diversity, and annotation quality of existing datasets. To address this, we introduce M3DLayout, a large-scale, multi-source dataset for 3D indoor layout generation. M3DLayout comprises 21,367 layouts and over 433k object instances, integrating three distinct sources: real-world scans, professional CAD designs, and procedurally generated scenes. Each layout is paired with detailed structured text describing global scene summaries, relational placements of large furniture, and fine-grained arrangements of smaller items. This diverse and richly annotated resource enables models to learn complex spatial and semantic patterns across a wide variety of indoor environments. To assess the potential of M3DLayout, we establish a benchmark using both a text-conditioned diffusion model and a text-conditioned autoregressive model. Experimental results demonstrate that our dataset provides a solid foundation for training layout generation models. Its multi-source composition enhances diversity, notably through the Inf3DLayout subset which provides rich small-object information, enabling the generation of more complex and detailed scenes. We hope that M3DLayout can serve as a valuable resource for advancing research in text-driven 3D scene synthesis. All dataset and code will be made public upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。