用图文输入生成精准三维建模,解决空间定位不准问题。
CAD-GPT: Synthesising CAD Construction Sequence with Spatial Reasoning-Enhanced Multimodal LLMs
- 将3D空间位置映射到一维语言空间,实现精确坐标推理。
- 在多个数据集上优于现有方法,生成模型更符合真实空间结构。
- 适合需要高精度建模的工程设计人员使用。
计算机辅助设计(CAD)通过精确的二维和三维建模、分析与优化,显著提升设计效率、准确性和创新性。现有建模方法依赖难以获取的潜在向量或点云,存储成本高。近年来,多模态大语言模型(MLLMs)被用于根据自然语言指令和图像生成CAD模型,但这些模型在推断三维空间位置与朝向方面仍存在困难,导致几何构建的起始点和拉伸方向不准确。本文提出CAD-GPT,一种基于空间推理增强的MLLM的CAD合成方法,可接受单张图像或文本描述作为输入。为实现精确的空间推断,我们引入3D建模空间机制,利用专门的空间展开机制将3D位置和草图平面旋转角映射至一维语言特征空间,同时将2D草图坐标离散化至适当平面空间,以精确定位起始点、草图朝向及坐标变换。大量实验表明,CAD-GPT在定量和定性评估中均持续优于现有最先进方法。
原文摘要 · Abstract (English)
Computer-aided design (CAD) significantly enhances the efficiency, accuracy, and innovation of design processes by enabling precise 2D and 3D modeling, extensive analysis, and optimization. Existing methods for creating CAD models rely on latent vectors or point clouds, which are difficult to obtain, and storage costs are substantial. Recent advances in Multimodal Large Language Models (MLLMs) have inspired researchers to use natural language instructions and images for CAD model construction. However, these models still struggle with inferring accurate 3D spatial location and orientation, leading to inaccuracies in determining the spatial 3D starting points and extrusion directions for constructing geometries. This work introduces CAD-GPT, a CAD synthesis method with spatial reasoning-enhanced MLLM that takes either a single image or a textual description as input. To achieve precise spatial inference, our approach introduces a 3D Modeling Spatial Mechanism. This method maps 3D spatial positions and 3D sketch plane rotation angles into a 1D linguistic feature space using a specialized spatial unfolding mechanism, while discretizing 2D sketch coordinates into an appropriate planar space to enable precise determination of spatial starting position, sketch orientation, and 2D sketch coordinate translations. Extensive experiments demonstrate that CAD-GPT consistently outperforms existing state-of-the-art methods in CAD model synthesis, both quantitatively and qualitatively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。