arXiv:2503.15475cs.CV2025-03被引 17

用3D形状分词器构建通用3D智能基础模型,支持从文本生成3D场景

Cube: A Roblox View of 3D Intelligence

  • 设计3D形状分词器,将几何形状转为可计算的离散单元
  • 实现文本到3D物体、场景的生成,支持与大语言模型协同推理
  • 面向游戏开发和内容创作,适合想快速搭建3D世界的开发者

在文本、图像、音频和视频领域,基于海量数据训练的基础模型已展现出卓越的推理与生成能力。我们的目标是在Roblox上构建一个3D智能基础模型,支持开发者完成从生成3D物体与场景、角色骨骼绑定到编写对象行为脚本的全流程。本文提出该模型的三大设计要求,并展示首个实现步骤:以3D几何形状为核心数据类型,设计3D形状分词器。该分词方案可应用于文本到形状、形状到文本及文本到场景的生成任务。同时,我们展示了这些应用如何与现有大语言模型(LLMs)协作,实现场景分析与推理。最后,讨论了构建统一3D智能基础模型的未来路径。

原文摘要 · Abstract (English)

Foundation models trained on vast amounts of data have demonstrated remarkable reasoning and generation capabilities in the domains of text, images, audio and video. Our goal at Roblox is to build such a foundation model for 3D intelligence, a model that can support developers in producing all aspects of a Roblox experience, from generating 3D objects and scenes to rigging characters for animation to producing programmatic scripts describing object behaviors. We discuss three key design requirements for such a 3D foundation model and then present our first step towards building such a model. We expect that 3D geometric shapes will be a core data type and describe our solution for 3D shape tokenizer. We show how our tokenization scheme can be used in applications for text-to-shape generation, shape-to-text generation and text-to-scene generation. We demonstrate how these applications can collaborate with existing large language models (LLMs) to perform scene analysis and reasoning. We conclude with a discussion outlining our path to building a fully unified foundation model for 3D intelligence.

3D生成基础模型游戏开发形状分词

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。