arXiv:2504.09328cs.CVcs.LG2025-04被引 1

用文字生成3D家具并自动组装到房间,解决场景数据不足问题

Text To 3D Object Generation For Scalable Room Assembly

  • 结合文本生成图像与多视角扩散模型,生成高保真3D物体
  • 支持按需生成带家具的3D室内场景,提升数据多样性
  • 适合需要大量3D场景数据的研究者与工业应用

当前用于场景理解的机器学习模型(如深度估计、物体追踪)依赖于大规模高质量数据集,但真实世界数据获取困难。为此,我们提出一个端到端的合成数据生成系统,可创建可扩展、高质量且可定制的3D室内场景。通过整合并改进文本到图像及多视角扩散模型,并结合基于神经辐射场的网格重建技术,该系统能从文本提示生成高保真3D物体资产,并利用渲染工具将其嵌入预设的平面图中。通过引入新型损失函数和训练策略,系统支持按需生成场景,旨在缓解现有数据普遍由艺术家手工制作导致的稀缺问题。该方法推动了合成数据在应对机器学习训练瓶颈中的作用,有助于构建更鲁棒、泛化能力强的真实世界应用模型。

原文摘要 · Abstract (English)

Modern machine learning models for scene understanding, such as depth estimation and object tracking, rely on large, high-quality datasets that mimic real-world deployment scenarios. To address data scarcity, we propose an end-to-end system for synthetic data generation for scalable, high-quality, and customizable 3D indoor scenes. By integrating and adapting text-to-image and multi-view diffusion models with Neural Radiance Field-based meshing, this system generates highfidelity 3D object assets from text prompts and incorporates them into pre-defined floor plans using a rendering tool. By introducing novel loss functions and training strategies into existing methods, the system supports on-demand scene generation, aiming to alleviate the scarcity of current available data, generally manually crafted by artists. This system advances the role of synthetic data in addressing machine learning training limitations, enabling more robust and generalizable models for real-world applications.

3D生成文本生成合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。