arXiv:2507.06484cs.GRcs.CV2025-07

用AI自动生成高质量3D环境,提升模型空间推理能力

3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds

  • 将3D场景构建转化为序列决策问题,用视觉语言模型生成布局材质等
  • 自进化微调后生成的3D数据使模型性能接近真实数据训练结果
  • 适合需要大量3D仿真数据的研究者,如机器人、VR和游戏开发

尽管大规模预训练赋予模型语言与视觉推理能力,但其空间推理能力仍因缺乏真实3D世界数据而受限。虽然人类可通过3D图形技术创建沉浸式交互世界(如虚拟现实、游戏、机器人),但过程极为耗时。本文提出一种可扩展的高质量3D环境生成方法,将3D场景构建视为序列决策问题,使用视觉语言模型(VLM)作为策略,联合输出布局、材质、光照和资产等动作。提出的框架3D-Generalist通过自进化微调,使VLM生成更符合提示的3D环境。我们验证了3D-Generalist在生成可用于训练的基础模型的仿真环境上的有效性。进一步证明其在合成数据生成中的质量和可扩展性:在生成数据上预训练视觉基础模型后,经下游任务微调,其性能超越在人工精心制作的合成数据上预训练的模型,并接近使用规模大得多的真实数据所达到的效果。

原文摘要 · Abstract (English)

Despite large-scale pretraining endowing models with language and vision reasoning capabilities, improving their spatial reasoning capability remains challenging due to the lack of data grounded in the 3D world. While it is possible for humans to manually create immersive and interactive worlds through 3D graphics, as seen in applications such as VR, gaming, and robotics, this process remains highly labor-intensive. In this paper, we propose a scalable method for generating high-quality 3D environments that can serve as training data for foundation models. We recast 3D environment building as a sequential decision-making problem, employing Vision-Language-Models (VLMs) as policies that output actions to jointly craft a 3D environment's layout, materials, lighting, and assets. Our proposed framework, 3D-Generalist, trains VLMs to generate more prompt-aligned 3D environments via self-improvement fine-tuning. We demonstrate the effectiveness of 3D-Generalist and the proposed training strategy in generating simulation-ready 3D environments. Furthermore, we demonstrate its quality and scalability in synthetic data generation by pretraining a vision foundation model on the generated data. After fine-tuning the pre-trained model on downstream tasks, we show that it surpasses models pre-trained on meticulously human-crafted synthetic data and approaches results achieved with real data orders of magnitude larger.

3D生成自进化视觉语言模型仿真数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。