arXiv:2409.17457cs.CVcs.AI2024-09ECCV被引 39

用多模态大模型生成参数化CAD草图,首次实现语言与视觉融合设计。

CadVLM: Bridging Language and Vision in the Generation of Parametric CAD Sketches

  • 融合图像与草图序列,用预训练模型理解工程设计逻辑。
  • 在自动补全、约束生成等任务上表现优于现有方法。
  • 适合工业设计、智能制造领域研究者参考。

参数化计算机辅助设计(CAD)是现代机械设计的核心,但面临参数化草图建模精度不足和缺乏适用于机械设计的实用评估指标等问题。本文利用在自然语言处理和计算机视觉中取得成功的预训练基础模型,开发专用于CAD的生成模型,使其能够理解复杂几何结构和设计推理,显著提升CAD技术能力。提出CadVLM,一种端到端的视觉-语言模型,用于CAD生成。该方法通过适配预训练模型,有效操控工程草图,整合草图基元序列与草图图像信息。大量实验表明,在多个CAD草图生成任务(如自动补全、自动约束、图像条件生成)中均表现优越。据我们所知,这是首个成功将多模态大语言模型应用于参数化CAD生成的研究,标志着计算机辅助机械设计领域的重要突破。

原文摘要 · Abstract (English)

Parametric Computer-Aided Design (CAD) is central to contemporary mechanical design. However, it encounters challenges in achieving precise parametric sketch modeling and lacks practical evaluation metrics suitable for mechanical design. We harness the capabilities of pre-trained foundation models, renowned for their successes in natural language processing and computer vision, to develop generative models specifically for CAD. These models are adept at understanding complex geometries and design reasoning, a crucial advancement in CAD technology. In this paper, we propose CadVLM, an end-to-end vision language model for CAD generation. Our approach involves adapting pre-trained foundation models to manipulate engineering sketches effectively, integrating both sketch primitive sequences and sketch images. Extensive experiments demonstrate superior performance on multiple CAD sketch generation tasks such as CAD autocompletion, CAD autoconstraint, and image conditional generation. To our knowledge, this is the first instance of a multimodal Large Language Model (LLM) being successfully applied to parametric CAD generation, representing a pioneering step in the field of computer-aided mechanical design.

CAD生成多模态模型参数化设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。