arXiv:2511.13647cs.CV2025-11被引 6

让3D模型理解语言指令,一键生成结构化编辑方案。

Part-X-MLLM: Part-aware 3D Multimodal Large Language Model

  • 用程序语法统一处理3D任务,输出含部件框、语义和指令的序列
  • 在多个任务上达到顶尖性能,支持自然语言直接控制3D编辑
  • 适合需要3D生成与编辑的科研和工业用户

我们提出 Part-X-MLLM,一种原生3D多模态大语言模型,通过将多样化的3D任务建模为结构化可执行语法中的程序来统一处理。给定RGB点云和自然语言提示,该模型自回归生成单一连贯的标记序列,包含部件级边界框、语义描述和编辑命令。此结构化输出作为通用接口,驱动下游几何感知模块实现基于部件的生成与编辑。通过将符号规划与几何合成解耦,本方法允许任意兼容的几何引擎通过单一语言原生前端进行控制。我们采用双编码器架构预训练以解耦结构与语义,并在大规模部件为中心的数据集上进行指令微调。实验表明,该模型能生成高质量、结构化的计划,在基于场景问答、组合生成和局部编辑任务中实现领先性能,且仅需一个统一接口。

原文摘要 · Abstract (English)

We introduce Part-X-MLLM, a native 3D multimodal large language model that unifies diverse 3D tasks by formulating them as programs in a structured, executable grammar. Given an RGB point cloud and a natural language prompt, our model autoregressively generates a single, coherent token sequence encoding part-level bounding boxes, semantic descriptions, and edit commands. This structured output serves as a versatile interface to drive downstream geometry-aware modules for part-based generation and editing. By decoupling the symbolic planning from the geometric synthesis, our approach allows any compatible geometry engine to be controlled through a single, language-native frontend. We pre-train a dual-encoder architecture to disentangle structure from semantics and instruction-tune the model on a large-scale, part-centric dataset. Experiments demonstrate that our model excels at producing high-quality, structured plans, enabling state-of-the-art performance in grounded Q\&A, compositional generation, and localized editing through one unified interface. Project page: https://chunshi.wang/Part-X-MLLM/

3D生成多模态语言控制结构化输出

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。