arXiv:2411.04954cs.CV2024-11被引 91

首个支持多模态输入生成参数化CAD模型的系统

CAD-MLLM: Unifying Multimodality-Conditioned CAD Generation With MLLM

  • 用指令序列与大语言模型统一多模态特征空间
  • 在45万实例上训练,生成模型拓扑质量高
  • 适合工业设计、自动化建模场景使用

本文旨在构建一个统一的计算机辅助设计(CAD)生成系统,可基于文本、图像、点云或其组合输入,快速生成参数化CAD模型。为此,我们提出CAD-MLLM,首个能根据多模态输入生成参数化CAD模型的系统。该框架利用CAD模型的命令序列,结合先进大语言模型(LLM),对齐多模态数据与CAD向量表示之间的特征空间。为支持训练,我们设计了完整的数据构建与标注流程,为每个CAD模型配以文本描述、多视角图像、点云及命令序列,形成首个包含四类模态数据的多模态CAD数据集Omni-CAD,共约45万实例及其构建序列。为全面评估生成质量,我们引入拓扑质量与表面封闭度等新指标,超越传统重建精度评价。大量实验表明,CAD-MLLM显著优于现有条件生成方法,且对噪声和缺失点具有强鲁棒性。

原文摘要 · Abstract (English)

This paper aims to design a unified Computer-Aided Design (CAD) generation system that can easily generate CAD models based on the user's inputs in the form of textual description, images, point clouds, or even a combination of them. Towards this goal, we introduce the CAD-MLLM, the first system capable of generating parametric CAD models conditioned on the multimodal input. Specifically, within the CAD-MLLM framework, we leverage the command sequences of CAD models and then employ advanced large language models (LLMs) to align the feature space across these diverse multi-modalities data and CAD models' vectorized representations. To facilitate the model training, we design a comprehensive data construction and annotation pipeline that equips each CAD model with corresponding multimodal data. Our resulting dataset, named Omni-CAD, is the first multimodal CAD dataset that contains textual description, multi-view images, points, and command sequence for each CAD model. It contains approximately 450K instances and their CAD construction sequences. To thoroughly evaluate the quality of our generated CAD models, we go beyond current evaluation metrics that focus on reconstruction quality by introducing additional metrics that assess topology quality and surface enclosure extent. Extensive experimental results demonstrate that CAD-MLLM significantly outperforms existing conditional generative methods and remains highly robust to noises and missing points. The project page and more visualizations can be found at: https://cad-mllm.github.io/

CAD生成多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。