IL3D为大模型生成3D室内场景提供超大规模数据集
IL3D: A Large-Scale Indoor Layout Dataset for LLM-Driven 3D Scene Generation
- 构建27,816个房间布局与29,215个高保真3D物体资产
- 在多任务评估中,微调后模型性能优于其他数据集
- 支持点云、多视角图等10余种模态输出,适配智能体研究
本文提出IL3D,一个面向大语言模型驱动的3D场景生成的大规模数据集,旨在满足室内布局设计中对多样化、高质量训练数据的迫切需求。该数据集包含18种常见房型的27,816个室内布局,以及29,215个高保真3D物体资产,并配备实例级自然语言标注,以支持视觉-语言任务的鲁棒多模态学习。我们建立了严格的评估基准,实验表明,在IL3D上进行监督微调(SFT)的LLM显著提升泛化能力,性能超越在其他数据集上的微调结果。IL3D支持点云、3D边界框、多视角图像、深度图、法向图和语义掩码等多种模态数据导出,可无缝适配各类视觉任务。作为多功能、强鲁棒性的资源,IL3D为3D场景生成与具身智能研究提供了高保真场景数据,助力智能体环境感知任务发展。
原文摘要 · Abstract (English)
In this study, we present IL3D, a large-scale dataset meticulously designed for large language model (LLM)-driven 3D scene generation, addressing the pressing demand for diverse, high-quality training data in indoor layout design. Comprising 27,816 indoor layouts across 18 prevalent room types and a library of 29,215 high-fidelity 3D object assets, IL3D is enriched with instance-level natural language annotations to support robust multimodal learning for vision-language tasks. We establish rigorous benchmarks to evaluate LLM-driven scene generation. Experimental results show that supervised fine-tuning (SFT) of LLMs on IL3D significantly improves generalization and surpasses the performance of SFT on other datasets. IL3D offers flexible multimodal data export capabilities, including point clouds, 3D bounding boxes, multiview images, depth maps, normal maps, and semantic masks, enabling seamless adaptation to various visual tasks. As a versatile and robust resource, IL3D significantly advances research in 3D scene generation and embodied intelligence, by providing high-fidelity scene data to support environment perception tasks of embodied agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。