用AI Agent生成结构化、可交互的3D模型,让文字直接变立体资产。
ShapeCraft: LLM Agents for Structured, Textured and Interactive 3D Modeling

- 将3D建模拆解为可执行的任务图,让大模型精准理解空间关系。
- 生成的3D模型几何准确、带纹理,支持动画与用户修改。
- 适合需要快速原型设计的艺术家和游戏开发者使用。
从自然语言生成3D模型有望大幅减少专家手动建模工作量,并提升3D资产的可及性。然而,现有方法常生成无结构网格且交互性差,难以融入艺术创作流程。为此,我们提出将3D资产表示为形状程序,并引入ShapeCraft——一种用于文本到3D生成的多智能体框架。核心是图结构过程化形状(GPS)表示,能将复杂自然语言分解为子任务结构图,从而促进大模型对空间关系和语义形状细节的理解。具体而言,LLM智能体分层解析用户输入以初始化GPS,再迭代优化过程化建模与着色,最终生成结构化、带纹理且可交互的3D资产。定性和定量实验表明,相比现有基于LLM的智能体,ShapeCraft在生成几何准确、语义丰富的3D资产方面表现更优。我们还通过动画示例和用户自定义编辑展示了其多功能性,凸显其在更广泛交互应用中的潜力。
原文摘要 · Abstract (English)
3D generation from natural language offers significant potential to reduce expert manual modeling efforts and enhance accessibility to 3D assets. However, existing methods often yield unstructured meshes and exhibit poor interactivity, making them impractical for artistic workflows. To address these limitations, we represent 3D assets as shape programs and introduce ShapeCraft, a novel multi-agent framework for text-to-3D generation. At its core, we propose a Graph-based Procedural Shape (GPS) representation that decomposes complex natural language into a structured graph of sub-tasks, thereby facilitating accurate LLM comprehension and interpretation of spatial relationships and semantic shape details. Specifically, LLM agents hierarchically parse user input to initialize GPS, then iteratively refine procedural modeling and painting to produce structured, textured, and interactive 3D assets. Qualitative and quantitative experiments demonstrate ShapeCraft's superior performance in generating geometrically accurate and semantically rich 3D assets compared to existing LLM-based agents. We further show the versatility of ShapeCraft through examples of animated and user-customized editing, highlighting its potential for broader interactive applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。