arXiv:2509.16768cs.CV2025-09被引 1

用视觉语言模型生成零件级3D模型,支持可控分割与遮挡还原。

MMPart: Harnessing Multi-Modal Large Language Models for Part-Aware 3D Generation

  • 基于图像和用户描述生成提示,引导模型分部件生成
  • 通过多视角一致图像实现遮挡区域的合理补全
  • 适合需要精细编辑和语义理解的3D内容创作

生成式3D建模在VR/AR、元宇宙和机器人等领域快速发展,但现有方法通常将物体表示为封闭网格,缺乏结构信息,限制了编辑、动画和语义理解。零件级3D生成通过将物体分解为有意义的组件来解决此问题,但现有流程存在挑战:用户无法控制哪些部分被分离,且模型在隔离阶段难以合理想象被遮挡区域。本文提出MMPart框架,仅需单张图像即可生成零件感知的3D模型。首先利用视觉语言模型(VLM)根据输入图像和用户描述生成一组提示;随后,生成模型以初始图像和前一步提示作为监督,生成各部件的孤立图像,提示控制姿态并引导模型推断被遮挡部分;接着进入多视角生成阶段,生成多个视角的一致图像;最后,重建模型将每组多视角图像转换为3D模型。

原文摘要 · Abstract (English)

Generative 3D modeling has advanced rapidly, driven by applications in VR/AR, metaverse, and robotics. However, most methods represent the target object as a closed mesh devoid of any structural information, limiting editing, animation, and semantic understanding. Part-aware 3D generation addresses this problem by decomposing objects into meaningful components, but existing pipelines face challenges: in existing methods, the user has no control over which objects are separated and how model imagine the occluded parts in isolation phase. In this paper, we introduce MMPart, an innovative framework for generating part-aware 3D models from a single image. We first use a VLM to generate a set of prompts based on the input image and user descriptions. In the next step, a generative model generates isolated images of each object based on the initial image and the previous step's prompts as supervisor (which control the pose and guide model how imagine previously occluded areas). Each of those images then enters the multi-view generation stage, where a number of consistent images from different views are generated. Finally, a reconstruction model converts each of these multi-view images into a 3D model.

3D生成多模态零件感知视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。