首个原生支持3D生成与理解的多模态大模型,可自由处理3D、文本混合输入。
ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding

- 用3D-VQVAE将物体编码为离散向量,实现高效3D表征
- 构建3D-Alpaca数据集,覆盖生成、理解、编辑全链条任务
- 基于Qwen微调,首次实现3D与文本任意顺序交互
近期,ChatGPT-4o展现出强大的图文生成能力,推动了原生多模态大语言模型的发展。然而其多模态能力仍局限于图像与文本。而对3D内容的理解与生成同样关键。为此,我们提出ShapeLLM-Omni——一个原生支持3D生成与理解的多模态大模型,可任意顺序处理3D资产、文本。首先,训练一个3D向量量化变分自编码器(VQVAE),将3D对象映射到离散潜在空间,实现高效精准的形状表示与重建。在此基础上,创新构建大规模连续训练数据集3D-Alpaca,涵盖生成、理解与编辑任务,为未来研究提供丰富资源。最后,在3D-Alpaca数据集上对Qwen-2.5-vl-7B-Instruct模型进行指令微调。本工作为扩展多模态模型至基础3D能力提供了有效尝试,助力未来3D原生AI研究。
原文摘要 · Abstract (English)
Recently, the powerful text-to-image capabilities of ChatGPT-4o have led to growing appreciation for native multimodal large language models. However, its multimodal capabilities remain confined to images and text. Yet beyond images, the ability to understand and generate 3D content is equally crucial. To address this gap, we propose ShapeLLM-Omni-a native 3D large language model capable of understanding and generating 3D assets and text in any sequence. First, we train a 3D vector-quantized variational autoencoder (VQVAE), which maps 3D objects into a discrete latent space to achieve efficient and accurate shape representation and reconstruction. Building upon the 3D-aware discrete tokens, we innovatively construct a large-scale continuous training dataset named 3D-Alpaca, encompassing generation, comprehension, and editing, thus providing rich resources for future research and training. Finally, by performing instruction-based training of the Qwen-2.5-vl-7B-Instruct model on the 3D-Alpaca dataset. Our work provides an effective attempt at extending multimodal models with basic 3D capabilities, which contributes to future research in 3D-native AI. Project page: https://github.com/JAMESYJL/ShapeLLM-Omni
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。