arXiv:2608.20448cs.GRcs.CV2026-08

让3D物体按零件语义和位置精准生成,支持专业设计流程

MultiCube: Compositional 3D Generation With Part-Level Semantic and Spatial Control

论文配图:MultiCube: Compositional 3D Generation With Part-Level Semantic and Spatial Control
图 1 · 摘自论文原文
  • 分两阶段生成:先构型再拆解,每部分独立控制
  • 输入文本提示+零件清单+边界框布局,输出带分块网格的3D物
  • 支持复杂布局生成,适合游戏动画等高精度创作场景

游戏与动画中的数字3D对象通常需为可组合结构,即分解为语义明确的部件。现有3D生成方法虽能基于图像或文本生成高质量组合对象,但全局条件控制难以实现精细的部件级调控,不满足专业创作需求。为此,本文提出MultiCube,一种新型组合式3D生成方法,可对每个部件的语义与空间布局进行显式、独立控制。该方法输入包括全局文本提示、指定部件的文本模板及各部件在模板中的边界框布局,输出为多个符合语义与空间约束的独立网格组成的3D物体。其采用两阶段扩散过程:先生成与模板和布局对齐的单一整体网格,再同步分解为各部件。引入新颖的部件布局适配器,使每部分条件独立编码。实验表明,该方法可在部件级实现精确控制,生成仅靠文本或图像提示难以实现的独特布局,显著提升生成质量与可控性。

原文摘要 · Abstract (English)

Digital 3D objects used in games and animation are often required to be compositional; that is, decomposed into semantically meaningful parts. Recent 3D generation methods can produce high-quality compositional objects conditioned on image or text prompts. Yet, such global conditioning lacks the precise part-level controllability required for professional creative workflows. To address this, we introduce MultiCube, a novel compositional 3D generation method that provides explicit, independent control over both the semantics and spatial arrangement of each part. MultiCube takes as input a global text prompt, a text schema specifying the desired parts, and a spatial layout indicating the bounding boxes of the parts in the given schema. It outputs a 3D object composed of distinct meshes, one per specified part, that adhere to the given semantic and spatial conditions. Our approach employs a two-stage diffusion process, first generating a schema- and layout-aligned monolithic mesh, then decomposing the mesh into individual parts simultaneously. A novel Part Layout Adapter is used to encode per-part conditions independently of the other parts. Experiments demonstrate that our method can generate high-quality compositional 3D objects with precise part-level control, including those with unique layouts difficult to achieve with text or image prompting alone. Project page: https://multi-cube.github.io

3D生成组合建模可控生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。